Step 1: Decide where recommendations appear, and what success means
Recommendations show up in a few typical places, and each one is a slightly different problem:
- “For you” on the home page or in emails: personal, based on everything the customer has bought.
- “Goes well with” on a product page: based on the product, so it works even for visitors you don’t know.
- In the basket: things often bought together with what’s already in it.
Before building anything, decide how you’ll measure it. Offline, use the hit rate: hide each customer’s most recent month, and check how many of them went on to buy at least one of the 10 products you would have shown. Online, the real test is whether people click and buy more, which only an A/B test can tell you (Step 8).
Step 2: Start from purchase history
You need one row per purchase: customer_id, product_id and date, plus a products table with at least a category. Export them from your shop or database. Views, clicks and ratings can help later, but purchases are the clearest signal of what people want.
No data to hand? This script makes a realistic sample for an online shop: 3,000 customers who each prefer two of eight categories, 400 products with a few best-sellers and a long tail, and pairs of products that are often bought together.
CodeOptional: make a sample purchases.csv and products.csv
# pip install pandas numpy
import numpy as np
import pandas as pd
rng = np.random.default_rng(11)
CATEGORIES = ["coffee", "tea", "baking", "kitchen", "garden", "pets", "kids", "outdoor"]
n_products, n_customers = 400, 3000
# Products: a category, and a popularity that follows a long tail (a few best-sellers, many niche items).
products = pd.DataFrame({
"product_id": np.arange(n_products),
"category": rng.choice(CATEGORIES, n_products),
"popularity": rng.pareto(1.5, n_products) + 1,
})
# Some products go together (a grinder with coffee beans): each has a partner in the same category.
products["partner"] = [rng.choice(products.index[products.category == c]) for c in products.category]
rows = []
for customer in range(n_customers):
likes = rng.choice(CATEGORIES, size=2, replace=False) # each customer cares about two categories
weight = products.popularity * np.where(products.category.isin(likes), 12, 1)
n_orders = rng.integers(2, 25)
days = np.sort(rng.integers(0, 365, n_orders))
for day in days:
item = rng.choice(n_products, p=weight / weight.sum())
rows.append((customer, item, day))
if rng.random() < 0.3: # often bought together
rows.append((customer, products.partner[item], day))
purchases = pd.DataFrame(rows, columns=["customer_id", "product_id", "day"])
purchases["date"] = pd.Timestamp("2025-01-01") + pd.to_timedelta(purchases.pop("day"), unit="D")
purchases.to_csv("purchases.csv", index=False)
products[["product_id", "category"]].to_csv("products.csv", index=False)
print(len(purchases), "purchases by", purchases.customer_id.nunique(), "customers")Step 3: Hold out the last month, honestly
Never test on random purchases: that lets the model peek at the future. Instead, learn from everything up to a cutoff date and test on what each customer bought in the following 30 days. Only count products that were new to them, because recommending something they already buy every week is easy and not very useful.
CodeSplit by time
import pandas as pd
purchases = pd.read_csv("purchases.csv", parse_dates=["date"]) # customer_id, product_id, date
products = pd.read_csv("products.csv") # product_id, category
def holdout(purchases, cutoff, days=30):
"""Learn from everything up to the cutoff; test on what each customer bought in the next 30 days
that they had never bought before."""
train = purchases[purchases.date <= cutoff]
later = purchases[(purchases.date > cutoff) & (purchases.date <= cutoff + pd.Timedelta(days=days))]
owned = train.groupby("customer_id").product_id.apply(set)
later = later[later.customer_id.isin(owned.index)]
new = later[[p not in owned[c] for c, p in zip(later.customer_id, later.product_id)]]
return train, owned, new.groupby("customer_id").product_id.apply(set)
test_cutoff = purchases.date.max() - pd.Timedelta(days=30)
train, owned, truth = holdout(purchases, test_cutoff)
print(f"Learning from {len(train)} purchases; testing on {len(truth)} customers who bought something new")Step 4: Measure the best-sellers list first
The baseline is the simplest possible recommendation: the most popular products the customer doesn’t already have. It’s what many shops show, and it’s surprisingly hard to beat, because popular products are popular for a reason.
CodeThe baseline and the hit rate
K = 10 # how many products we show
def hit_rate(recommend, truth):
"""Share of customers who bought at least one of the K products we would have shown them."""
return sum(len(set(recommend(c)) & bought) > 0 for c, bought in truth.items()) / len(truth)
def most_popular(train, owned):
"""The baseline: best-sellers the customer doesn't already have."""
popular = train.product_id.value_counts().index.tolist()
return lambda c: [p for p in popular if p not in owned[c]][:K]
print(f"Most popular: {hit_rate(most_popular(train, owned), truth):.1%}")Step 5: Customers who bought this also bought
Collaborative filtering finds products that the same customers tend to buy. Two products count as similar when many customers bought both. Each customer then gets the products most similar to everything they bought before. It needs no product descriptions, and every recommendation can be explained: “because you bought…”.
CodeItem-to-item similarity
# pip install scikit-learn scipy
import numpy as np
from scipy.sparse import csr_matrix
from sklearn.preprocessing import normalize
def purchase_matrix(train):
"""One row per customer, one column per product: 1 if they ever bought it."""
customers = train.customer_id.unique()
row = {c: i for i, c in enumerate(customers)}
pairs = train.drop_duplicates(["customer_id", "product_id"])
matrix = csr_matrix((np.ones(len(pairs)), (pairs.customer_id.map(row), pairs.product_id)),
shape=(len(customers), len(products)))
return matrix, row
def top_k(scores, owned_items):
scores = np.asarray(scores, dtype=float).ravel()
scores[list(owned_items)] = -np.inf # never recommend what they already have
return np.argsort(-scores)[:K].tolist()
def also_bought(train, owned):
"""Products are similar when the same customers buy both; score each product by how similar
it is to everything this customer bought."""
matrix, row = purchase_matrix(train)
item_sim = (normalize(matrix.T) @ normalize(matrix.T).T).toarray()
np.fill_diagonal(item_sim, 0)
return lambda c: top_k(matrix[row[c]] @ item_sim, owned[c])
print(f"Customers who bought this also bought: {hit_rate(also_bought(train, owned), truth):.1%}")On the sample data, this reaches about 34% of customers, against about 21% for the best-sellers list: roughly 60% better, from a few lines of code.
Step 6: Try a personalized model, and tune it on a separate month
Matrix factorization learns a few “taste” numbers for every customer and product, and recommends the products closest to each customer’s taste. The number of tastes matters a lot, so choose it on the month before the test month. Choosing it on the test month itself would make the result look better than it really is.
CodeMatrix factorization, tuned honestly
from sklearn.decomposition import TruncatedSVD
def for_you(train, owned, tastes):
"""Matrix factorization: describe each customer and product with a few "taste" numbers,
learned so that customers sit close to the products they buy."""
matrix, row = purchase_matrix(train)
svd = TruncatedSVD(n_components=tastes, random_state=0)
customer_taste = svd.fit_transform(matrix)
return lambda c: top_k(customer_taste[row[c]] @ svd.components_, owned[c])
# Choose how many tastes on the month *before* the test month, so the test stays honest.
val_train, val_owned, val_truth = holdout(train, test_cutoff - pd.Timedelta(days=30))
scores = {t: hit_rate(for_you(val_train, val_owned, t), val_truth) for t in [4, 8, 16, 32, 64]}
print("Validation month:", {t: f"{s:.1%}" for t, s in scores.items()})
best = max(scores, key=scores.get)
print(f"Personalized with {best} tastes, on the test month: {hit_rate(for_you(train, owned, best), truth):.1%}")Here the personalized model scores about 35%, practically the same as “also bought”. That’s a common result. When two approaches are this close, choose the simpler one: it’s easier to explain, to debug and to keep running. Switch only when a more complex model wins clearly on your own data.
Step 7: Product pages and new customers
On a product page, show the products most similar to that product. It works for every visitor, including people who aren’t logged in. A brand-new customer has no history yet, which is the cold start problem: show best-sellers until they’ve bought something, or best-sellers from the category they’re browsing.
Code“Goes well with” for a product page
def goes_with(train):
"""For a product page there's no customer to personalize for: show the products most often
bought by the same people."""
matrix, _ = purchase_matrix(train)
item_sim = (normalize(matrix.T) @ normalize(matrix.T).T).toarray()
np.fill_diagonal(item_sim, 0)
return lambda product: np.argsort(-item_sim[product])[:K].tolist()
similar = goes_with(train)
example = train.product_id.value_counts().index[0]
shown = products.set_index("product_id").loc[similar(example), "category"]
print(f"Product {example} ({products.category[example]}): suggestions are",
f"{(shown == products.category[example]).mean():.0%} from the same category")Step 8: Refresh every night, and prove it works
Recommendations don’t need to be computed while the page loads. A nightly batch job learns from all purchases so far and saves the top 10 for every customer; your shop reads them from a table. Every cloud has a cheap scheduled-job service; see the cloud comparison.
CodeNightly recommendations for every customer
# Every night: learn from all purchases so far, and save the top 10 for every customer.
owned_now = purchases.groupby("customer_id").product_id.apply(set)
recommend = also_bought(purchases, owned_now)
best_sellers = purchases.product_id.value_counts().index[:K].tolist()
def for_customer(customer):
# New customers have no history yet: show best-sellers until they've bought something.
return recommend(customer) if customer in owned_now.index else best_sellers
rows = [(c, rank + 1, p) for c in owned_now.index for rank, p in enumerate(for_customer(c))]
pd.DataFrame(rows, columns=["customer_id", "rank", "product_id"]).to_csv("recommendations.csv", index=False)
print(f"Saved {len(rows)} recommendations for {len(owned_now)} customers; a new customer sees {for_customer(-1)[:3]}…")Before switching everyone over, run an A/B test: show half of your customers the new recommendations and the other half the best-sellers list, and compare clicks and sales over a few weeks. Offline hit rates tell you which approach to try; only the A/B test tells you whether it actually earns more.
Common pitfalls
- Testing on random purchases instead of a later period, which lets the model see the future.
- Tuning the model on the test month, so the score looks better than it will be in real life.
- Recommending products people already own, or ones that are out of stock.
- Only ever showing best-sellers, so niche products never get a chance to be discovered.
- Trusting offline numbers alone: they pick the approach to test, but only an A/B test proves the value.