Recommendation System
Suggest relevant products, content or people to each user.
Typical projects
Three ways to build it
Starter
Popularity + content-based similarity (embeddings of item descriptions)Works with little interaction data and handles new items.
Best for: No or little data, or new to ML
sentence-transformers · pandas · scikit-learn
Standard
Collaborative filtering (implicit ALS / LightFM) + content featuresLearns taste from behaviour; proven and cheap to run.
Best for: Some labeled data and Python experience
implicit · LightFM · pandas
Advanced
Two-tower retrieval + ranking model (+ real-time features)Industry-standard architecture for millions of users and items.
Best for: Lots of data and an experienced team
TensorFlow Recommenders / TorchRec · FAISS · Feast · LightGBM ranker
How success is measured
Recall@K / NDCG offline; click-through and conversion in an A/B test online
The data you'll need
- Interaction logs are the key: user_id, item_id, timestamp, event (view/click/buy).
- Item metadata (title, category, description) helps with new items (cold start).
- Thousands of users with several interactions each is a good start.
Labeling
Implicit feedback (clicks, purchases) acts as labels — no manual labeling.
Preparing the data
- Deduplicate events, remove bots
- Split by time: train on the past, test on the next period
- Build item text/feature representations for cold start
Start with a baseline
Recommend the most popular items (overall or per category). Many systems barely beat this — measure it!
Evaluating the model
- Recall@10 on a time-based holdout
- Coverage & diversity (are you only showing best-sellers?)
- Online A/B test vs popularity baseline
Monitoring in production
- CTR and conversion
- Catalog coverage
- Feedback loops (popular gets more popular)
Common pitfalls
- Random splits leak future behaviour
- Ignoring the cold-start problem for new users/items
Example code
CodeQuick start
# Content-based "similar items"
# pip install sentence-transformers pandas
import pandas as pd
from sentence_transformers import SentenceTransformer, util
items = pd.read_csv("items.csv") # id, title, description
model = SentenceTransformer("all-MiniLM-L6-v2")
emb = model.encode((items.title + ". " + items.description).tolist(), convert_to_tensor=True)
scores = util.cos_sim(emb[0], emb)[0]
print(items.iloc[scores.argsort(descending=True)[1:6].cpu()].title)CodeTrain your own model
# pip install implicit pandas scipy
import pandas as pd, scipy.sparse as sp
from implicit.als import AlternatingLeastSquares
ev = pd.read_csv("events.csv") # user_id, item_id, weight
users = ev.user_id.astype("category"); items = ev.item_id.astype("category")
mat = sp.csr_matrix((ev.weight, (users.cat.codes, items.cat.codes)))
model = AlternatingLeastSquares(factors=64, iterations=20)
model.fit(mat)
uid = 0
ids, scores = model.recommend(uid, mat[uid], N=10)
print([items.cat.categories[i] for i in ids])