Hello Model
← Guides

Tabular Classification · 4 min read

How to predict customer churn from your order history

Use the order history you already have to rank customers by how likely they are to stop buying, so your team can reach out to the right people before they leave.

What you'll build

  • A clear, measurable definition of churn for your business
  • Features built from order history without leaking the future
  • A model that ranks customers by churn risk, refreshed every week

Step 1: Define churn precisely

"Churn" has to become a yes/no question you can answer from data. For a subscription business it's simple: did the customer cancel? For a shop without subscriptions, use a time window:

Churned = an active customer who places no order in the 60 days after a given date (the cutoff).

Pick the window from how often your customers normally buy. If most buy monthly, 60 days of silence is meaningful; if they buy twice a year, you'd need a longer window.

Also decide what you'll do with the predictions, because that decides how to measure success. A common plan: each week, the retention team contacts the 10% of customers with the highest risk. So the number that matters is how many of those top 10% really would have churned, alongside ROC-AUC for overall ranking quality.

Step 2: Start from an orders table

All you need is one row per order with three columns: customer_id, order_date and amount. Export it from your shop or database as a CSV. Customer details, support tickets or website visits can be added later as extra features.

Step 3: Build features as of a cutoff date

This is the most important idea in the guide. For each customer, compute features using only orders up to the cutoff, and the label using only orders after it. Mixing the two is data leakage: the model looks brilliant in testing and fails in real use, because in real life you never know the future.

CodePoint-in-time features and labels
python
# pip install pandas lightgbm scikit-learn shap
import pandas as pd

orders = pd.read_csv("orders.csv", parse_dates=["order_date"])  # customer_id, order_date, amount

def features(orders, cutoff, active_days=180):
    """What we knew about each active customer on the cutoff date."""
    past = orders[orders.order_date <= cutoff]
    f = past.groupby("customer_id").agg(
        last_order=("order_date", "max"),
        first_order=("order_date", "min"),
        orders=("order_date", "count"),
        total_spent=("amount", "sum"),
        avg_order=("amount", "mean"),
    )
    f["recency_days"] = (cutoff - f.pop("last_order")).dt.days
    f["tenure_days"] = (cutoff - f.pop("first_order")).dt.days
    recent = past[past.order_date > cutoff - pd.Timedelta(days=90)]
    f["orders_last_90d"] = recent.groupby("customer_id").size().reindex(f.index, fill_value=0)
    return f[f.recency_days <= active_days]   # long-gone customers aren't "at risk", they've left

def labels(orders, customers, cutoff, horizon_days=60):
    """1 if the customer placed no order in the horizon after the cutoff."""
    window = orders[(orders.order_date > cutoff) &
                    (orders.order_date <= cutoff + pd.Timedelta(days=horizon_days))]
    return pd.Series(~customers.isin(window.customer_id), index=customers, name="churned").astype(int)

Step 4: Train on the past, test on the more recent past

Instead of a random split, build the training set at one cutoff and the test set at a later one. That's exactly how the model will be used: trained on history, then applied to the next period. Make sure the training label window ends before the test cutoff.

CodeTwo cutoffs, no overlap
python
train_cut = pd.Timestamp("2026-03-31")   # labels use April–May
test_cut = pd.Timestamp("2026-06-30")    # labels use July–August

X_train = features(orders, train_cut)
y_train = labels(orders, X_train.index, train_cut)
X_test = features(orders, test_cut)
y_test = labels(orders, X_test.index, test_cut)
print(f"Churn rate: {y_train.mean():.0%} train, {y_test.mean():.0%} test")
Tip Your order history needs to run at least 60 days past the test cutoff, so the test labels are complete. Pick cutoffs that fit your data.

Step 5: Measure a simple baseline

Before training anything, check how well a single obvious rule ranks customers: "the longer since their last order, the more likely they've churned". Any model has to beat this baseline to be worth it.

CodeRecency as the baseline
python
from sklearn.metrics import roc_auc_score

print("Baseline ROC-AUC (recency only):", round(roc_auc_score(y_test, X_test.recency_days), 3))

Step 6: Train a gradient-boosted model

Gradient-boosted trees such as LightGBM are the go-to choice for tabular data: accurate, fast on a normal CPU, and able to handle features on very different scales without preparation. Then check the number that matters for the retention team.

CodeTrain LightGBM and check the top 10%
python
from lightgbm import LGBMClassifier

model = LGBMClassifier(n_estimators=300, learning_rate=0.05, num_leaves=15, min_child_samples=50, verbose=-1)
model.fit(X_train, y_train)

risk = model.predict_proba(X_test)[:, 1]
print("Model ROC-AUC:", round(roc_auc_score(y_test, risk), 3))

top = pd.Series(risk, index=X_test.index).nlargest(len(risk) // 10).index
print(f"Churn rate in the riskiest 10%: {y_test[top].mean():.0%} (overall: {y_test.mean():.0%})")

If the riskiest 10% churn at, say, three times the overall rate, the team's outreach is three times better targeted than contacting customers at random. That's the result to share with the business, not just the ROC-AUC.

If the model doesn't beat the recency baseline, that's a useful result too: use the simple rule for now, and add richer features such as support tickets, returns, discounts used or website visits, which carry signals that order history alone doesn't.

Step 7: Explain what drives the risk

People act on predictions they understand. SHAP shows how much each feature pushed each customer's risk up or down.

CodeFeature effects with SHAP
python
import shap

explanation = shap.TreeExplainer(model)(X_test)
shap.plots.beeswarm(explanation)   # one dot per customer, per feature

Expect recency and recent order counts to dominate. If a feature you didn't expect is at the top, check it for leakage before celebrating.

Step 8: Score customers every week

Before going live, retrain on the most recent complete window (cutoff = 60 days ago), so the model learns from the latest behaviour. Then, each week, compute features with today as the cutoff, with no labels needed, and hand the list to the team.

CodeWeekly scoring
python
today = pd.Timestamp.today().normalize()
X_now = features(orders, today)
scores = pd.Series(model.predict_proba(X_now)[:, 1], index=X_now.index, name="churn_risk")
scores.nlargest(len(scores) // 10).to_csv("at_risk_customers.csv")

A scheduled job that runs this script weekly is all the infrastructure you need. There's no need for a real-time API. Every cloud has a cheap batch or scheduled-job service; see the cloud comparison.

Step 9: Prove it works with a control group

A good model doesn't automatically mean fewer customers leave. Only the retention action can do that. To measure it, randomly keep part of the high-risk list (for example 20%) out of the campaign. Compare how many churn in each group after 60 days. This A/B test is the only honest way to show the project's value.

Common pitfalls

  • Building features with data from after the cutoff, for example "total orders this year" computed today.
  • Including customers who left long ago. They inflate the churn rate and teach the model nothing useful.
  • Random train/test splits across time, which let the model peek at the future.
  • Measuring the model but never the campaign. Without a control group you can't tell if outreach helped.

Want this tailored to your data, team and budget? Get a personalised plan →

More on Tabular Classification · Next guide: How to build a spam filter, start to finish