Hello Model
← Guides

Recommendation system · 4 min read

How to recommend products your customers will like

Turn your order history into a list of products each customer is likely to buy next, and into “goes well with” suggestions on every product page.

What you’ll build

  • Personal recommendations for every customer, refreshed every night
  • “Goes well with” suggestions for every product page
  • An honest measure of how much better they are than a best-sellers list

Step 1: Decide where recommendations appear, and what success means

Recommendations show up in a few typical places, and each one is a slightly different problem:

  • “For you” on the home page or in emails: personal, based on everything the customer has bought.
  • “Goes well with” on a product page: based on the product, so it works even for visitors you don’t know.
  • In the basket: things often bought together with what’s already in it.

Before building anything, decide how you’ll measure it. Offline, use the hit rate: hide each customer’s most recent month, and check how many of them went on to buy at least one of the 10 products you would have shown. Online, the real test is whether people click and buy more, which only an A/B test can tell you (Step 8).

Step 2: Start from purchase history

You need one row per purchase: customer_id, product_id and date, plus a products table with at least a category. Export them from your shop or database. Views, clicks and ratings can help later, but purchases are the clearest signal of what people want.

No data to hand? This script makes a realistic sample for an online shop: 3,000 customers who each prefer two of eight categories, 400 products with a few best-sellers and a long tail, and pairs of products that are often bought together.

CodeOptional: make a sample purchases.csv and products.csv
pythonRun in Colab
# pip install pandas numpy
import numpy as np
import pandas as pd

rng = np.random.default_rng(11)
CATEGORIES = ["coffee", "tea", "baking", "kitchen", "garden", "pets", "kids", "outdoor"]
n_products, n_customers = 400, 3000

# Products: a category, and a popularity that follows a long tail (a few best-sellers, many niche items).
products = pd.DataFrame({
    "product_id": np.arange(n_products),
    "category": rng.choice(CATEGORIES, n_products),
    "popularity": rng.pareto(1.5, n_products) + 1,
})
# Some products go together (a grinder with coffee beans): each has a partner in the same category.
products["partner"] = [rng.choice(products.index[products.category == c]) for c in products.category]

rows = []
for customer in range(n_customers):
    likes = rng.choice(CATEGORIES, size=2, replace=False)          # each customer cares about two categories
    weight = products.popularity * np.where(products.category.isin(likes), 12, 1)
    n_orders = rng.integers(2, 25)
    days = np.sort(rng.integers(0, 365, n_orders))
    for day in days:
        item = rng.choice(n_products, p=weight / weight.sum())
        rows.append((customer, item, day))
        if rng.random() < 0.3:                                       # often bought together
            rows.append((customer, products.partner[item], day))

purchases = pd.DataFrame(rows, columns=["customer_id", "product_id", "day"])
purchases["date"] = pd.Timestamp("2025-01-01") + pd.to_timedelta(purchases.pop("day"), unit="D")
purchases.to_csv("purchases.csv", index=False)
products[["product_id", "category"]].to_csv("products.csv", index=False)
print(len(purchases), "purchases by", purchases.customer_id.nunique(), "customers")

Step 3: Hold out the last month, honestly

Never test on random purchases: that lets the model peek at the future. Instead, learn from everything up to a cutoff date and test on what each customer bought in the following 30 days. Only count products that were new to them, because recommending something they already buy every week is easy and not very useful.

CodeSplit by time
pythonRun in Colab
import pandas as pd

purchases = pd.read_csv("purchases.csv", parse_dates=["date"])   # customer_id, product_id, date
products = pd.read_csv("products.csv")                            # product_id, category

def holdout(purchases, cutoff, days=30):
    """Learn from everything up to the cutoff; test on what each customer bought in the next 30 days
    that they had never bought before."""
    train = purchases[purchases.date <= cutoff]
    later = purchases[(purchases.date > cutoff) & (purchases.date <= cutoff + pd.Timedelta(days=days))]
    owned = train.groupby("customer_id").product_id.apply(set)
    later = later[later.customer_id.isin(owned.index)]
    new = later[[p not in owned[c] for c, p in zip(later.customer_id, later.product_id)]]
    return train, owned, new.groupby("customer_id").product_id.apply(set)

test_cutoff = purchases.date.max() - pd.Timedelta(days=30)
train, owned, truth = holdout(purchases, test_cutoff)
print(f"Learning from {len(train)} purchases; testing on {len(truth)} customers who bought something new")

Step 4: Measure the best-sellers list first

The baseline is the simplest possible recommendation: the most popular products the customer doesn’t already have. It’s what many shops show, and it’s surprisingly hard to beat, because popular products are popular for a reason.

CodeThe baseline and the hit rate
pythonRun in Colab
K = 10   # how many products we show

def hit_rate(recommend, truth):
    """Share of customers who bought at least one of the K products we would have shown them."""
    return sum(len(set(recommend(c)) & bought) > 0 for c, bought in truth.items()) / len(truth)

def most_popular(train, owned):
    """The baseline: best-sellers the customer doesn't already have."""
    popular = train.product_id.value_counts().index.tolist()
    return lambda c: [p for p in popular if p not in owned[c]][:K]

print(f"Most popular: {hit_rate(most_popular(train, owned), truth):.1%}")

Step 5: Customers who bought this also bought

Collaborative filtering finds products that the same customers tend to buy. Two products count as similar when many customers bought both. Each customer then gets the products most similar to everything they bought before. It needs no product descriptions, and every recommendation can be explained: “because you bought…”.

CodeItem-to-item similarity
pythonRun in Colab
# pip install scikit-learn scipy
import numpy as np
from scipy.sparse import csr_matrix
from sklearn.preprocessing import normalize

def purchase_matrix(train):
    """One row per customer, one column per product: 1 if they ever bought it."""
    customers = train.customer_id.unique()
    row = {c: i for i, c in enumerate(customers)}
    pairs = train.drop_duplicates(["customer_id", "product_id"])
    matrix = csr_matrix((np.ones(len(pairs)), (pairs.customer_id.map(row), pairs.product_id)),
                        shape=(len(customers), len(products)))
    return matrix, row

def top_k(scores, owned_items):
    scores = np.asarray(scores, dtype=float).ravel()
    scores[list(owned_items)] = -np.inf                     # never recommend what they already have
    return np.argsort(-scores)[:K].tolist()

def also_bought(train, owned):
    """Products are similar when the same customers buy both; score each product by how similar
    it is to everything this customer bought."""
    matrix, row = purchase_matrix(train)
    item_sim = (normalize(matrix.T) @ normalize(matrix.T).T).toarray()
    np.fill_diagonal(item_sim, 0)
    return lambda c: top_k(matrix[row[c]] @ item_sim, owned[c])

print(f"Customers who bought this also bought: {hit_rate(also_bought(train, owned), truth):.1%}")

On the sample data, this reaches about 34% of customers, against about 21% for the best-sellers list: roughly 60% better, from a few lines of code.

Step 6: Try a personalized model, and tune it on a separate month

Matrix factorization learns a few “taste” numbers for every customer and product, and recommends the products closest to each customer’s taste. The number of tastes matters a lot, so choose it on the month before the test month. Choosing it on the test month itself would make the result look better than it really is.

CodeMatrix factorization, tuned honestly
pythonRun in Colab
from sklearn.decomposition import TruncatedSVD

def for_you(train, owned, tastes):
    """Matrix factorization: describe each customer and product with a few "taste" numbers,
    learned so that customers sit close to the products they buy."""
    matrix, row = purchase_matrix(train)
    svd = TruncatedSVD(n_components=tastes, random_state=0)
    customer_taste = svd.fit_transform(matrix)
    return lambda c: top_k(customer_taste[row[c]] @ svd.components_, owned[c])

# Choose how many tastes on the month *before* the test month, so the test stays honest.
val_train, val_owned, val_truth = holdout(train, test_cutoff - pd.Timedelta(days=30))
scores = {t: hit_rate(for_you(val_train, val_owned, t), val_truth) for t in [4, 8, 16, 32, 64]}
print("Validation month:", {t: f"{s:.1%}" for t, s in scores.items()})
best = max(scores, key=scores.get)
print(f"Personalized with {best} tastes, on the test month: {hit_rate(for_you(train, owned, best), truth):.1%}")

Here the personalized model scores about 35%, practically the same as “also bought”. That’s a common result. When two approaches are this close, choose the simpler one: it’s easier to explain, to debug and to keep running. Switch only when a more complex model wins clearly on your own data.

Tip With too many taste numbers, the model memorizes each customer’s past and does worse than the best-sellers list (64 tastes scored 15% in validation). With too few, it can’t tell tastes apart. The validation month is what shows you where the sweet spot is.

Step 7: Product pages and new customers

On a product page, show the products most similar to that product. It works for every visitor, including people who aren’t logged in. A brand-new customer has no history yet, which is the cold start problem: show best-sellers until they’ve bought something, or best-sellers from the category they’re browsing.

Code“Goes well with” for a product page
pythonRun in Colab
def goes_with(train):
    """For a product page there's no customer to personalize for: show the products most often
    bought by the same people."""
    matrix, _ = purchase_matrix(train)
    item_sim = (normalize(matrix.T) @ normalize(matrix.T).T).toarray()
    np.fill_diagonal(item_sim, 0)
    return lambda product: np.argsort(-item_sim[product])[:K].tolist()

similar = goes_with(train)
example = train.product_id.value_counts().index[0]
shown = products.set_index("product_id").loc[similar(example), "category"]
print(f"Product {example} ({products.category[example]}): suggestions are",
      f"{(shown == products.category[example]).mean():.0%} from the same category")

Step 8: Refresh every night, and prove it works

Recommendations don’t need to be computed while the page loads. A nightly batch job learns from all purchases so far and saves the top 10 for every customer; your shop reads them from a table. Every cloud has a cheap scheduled-job service; see the cloud comparison.

CodeNightly recommendations for every customer
pythonRun in Colab
# Every night: learn from all purchases so far, and save the top 10 for every customer.
owned_now = purchases.groupby("customer_id").product_id.apply(set)
recommend = also_bought(purchases, owned_now)
best_sellers = purchases.product_id.value_counts().index[:K].tolist()

def for_customer(customer):
    # New customers have no history yet: show best-sellers until they've bought something.
    return recommend(customer) if customer in owned_now.index else best_sellers

rows = [(c, rank + 1, p) for c in owned_now.index for rank, p in enumerate(for_customer(c))]
pd.DataFrame(rows, columns=["customer_id", "rank", "product_id"]).to_csv("recommendations.csv", index=False)
print(f"Saved {len(rows)} recommendations for {len(owned_now)} customers; a new customer sees {for_customer(-1)[:3]}…")

Before switching everyone over, run an A/B test: show half of your customers the new recommendations and the other half the best-sellers list, and compare clicks and sales over a few weeks. Offline hit rates tell you which approach to try; only the A/B test tells you whether it actually earns more.

Common pitfalls

  • Testing on random purchases instead of a later period, which lets the model see the future.
  • Tuning the model on the test month, so the score looks better than it will be in real life.
  • Recommending products people already own, or ones that are out of stock.
  • Only ever showing best-sellers, so niche products never get a chance to be discovered.
  • Trusting offline numbers alone: they pick the approach to test, but only an A/B test proves the value.

Want this tailored to your data, team and budget? Create a plan for your project

More on Recommendation system · Next guide: How to build a spam filter, start to finish