Hello Model
← Model library

Recommendation System

Suggest relevant products, content or people to each user.

Typical projects

Recommend products based on purchase history'You might also like' for articlesSuggest courses to learners

Three ways to build it

Starter

Popularity + content-based similarity (embeddings of item descriptions)

Works with little interaction data and handles new items.

Best for: No or little data, or new to ML

sentence-transformers · pandas · scikit-learn

Standard

Collaborative filtering (implicit ALS / LightFM) + content features

Learns taste from behaviour; proven and cheap to run.

Best for: Some labeled data and Python experience

implicit · LightFM · pandas

Advanced

Two-tower retrieval + ranking model (+ real-time features)

Industry-standard architecture for millions of users and items.

Best for: Lots of data and an experienced team

TensorFlow Recommenders / TorchRec · FAISS · Feast · LightGBM ranker

How success is measured

Recall@K / NDCG offline; click-through and conversion in an A/B test online

The data you'll need

  • Interaction logs are the key: user_id, item_id, timestamp, event (view/click/buy).
  • Item metadata (title, category, description) helps with new items (cold start).
  • Thousands of users with several interactions each is a good start.

Labeling

Implicit feedback (clicks, purchases) acts as labels — no manual labeling.

Preparing the data

  • Deduplicate events, remove bots
  • Split by time: train on the past, test on the next period
  • Build item text/feature representations for cold start

Start with a baseline

Recommend the most popular items (overall or per category). Many systems barely beat this — measure it!

Evaluating the model

  • Recall@10 on a time-based holdout
  • Coverage & diversity (are you only showing best-sellers?)
  • Online A/B test vs popularity baseline

Monitoring in production

  • CTR and conversion
  • Catalog coverage
  • Feedback loops (popular gets more popular)

Common pitfalls

  • Random splits leak future behaviour
  • Ignoring the cold-start problem for new users/items

Example code

CodeQuick start
python
# Content-based "similar items"
# pip install sentence-transformers pandas
import pandas as pd
from sentence_transformers import SentenceTransformer, util

items = pd.read_csv("items.csv")            # id, title, description
model = SentenceTransformer("all-MiniLM-L6-v2")
emb = model.encode((items.title + ". " + items.description).tolist(), convert_to_tensor=True)
scores = util.cos_sim(emb[0], emb)[0]
print(items.iloc[scores.argsort(descending=True)[1:6].cpu()].title)
CodeTrain your own model
python
# pip install implicit pandas scipy
import pandas as pd, scipy.sparse as sp
from implicit.als import AlternatingLeastSquares

ev = pd.read_csv("events.csv")              # user_id, item_id, weight
users = ev.user_id.astype("category"); items = ev.item_id.astype("category")
mat = sp.csr_matrix((ev.weight, (users.cat.codes, items.cat.codes)))

model = AlternatingLeastSquares(factors=64, iterations=20)
model.fit(mat)
uid = 0
ids, scores = model.recommend(uid, mat[uid], N=10)
print([items.cat.categories[i] for i in ids])

Ready to build one? Get a personalised plan →

Or read about Anomaly Detection next.