Step 1: Define churn precisely
"Churn" has to become a yes/no question you can answer from data. For a subscription business it's simple: did the customer cancel? For a shop without subscriptions, use a time window:
Pick the window from how often your customers normally buy. If most buy monthly, 60 days of silence is meaningful; if they buy twice a year, you'd need a longer window.
Also decide what you'll do with the predictions, because that decides how to measure success. A common plan: each week, the retention team contacts the 10% of customers with the highest risk. So the number that matters is how many of those top 10% really would have churned, alongside ROC-AUC for overall ranking quality.
Step 2: Start from an orders table
All you need is one row per order with three columns: customer_id, order_date and amount. Export it from your shop or database as a CSV. Customer details, support tickets or website visits can be added later as extra features.
Step 3: Build features as of a cutoff date
This is the most important idea in the guide. For each customer, compute features using only orders up to the cutoff, and the label using only orders after it. Mixing the two is data leakage: the model looks brilliant in testing and fails in real use, because in real life you never know the future.
CodePoint-in-time features and labels
# pip install pandas lightgbm scikit-learn shap
import pandas as pd
orders = pd.read_csv("orders.csv", parse_dates=["order_date"]) # customer_id, order_date, amount
def features(orders, cutoff, active_days=180):
"""What we knew about each active customer on the cutoff date."""
past = orders[orders.order_date <= cutoff]
f = past.groupby("customer_id").agg(
last_order=("order_date", "max"),
first_order=("order_date", "min"),
orders=("order_date", "count"),
total_spent=("amount", "sum"),
avg_order=("amount", "mean"),
)
f["recency_days"] = (cutoff - f.pop("last_order")).dt.days
f["tenure_days"] = (cutoff - f.pop("first_order")).dt.days
recent = past[past.order_date > cutoff - pd.Timedelta(days=90)]
f["orders_last_90d"] = recent.groupby("customer_id").size().reindex(f.index, fill_value=0)
return f[f.recency_days <= active_days] # long-gone customers aren't "at risk", they've left
def labels(orders, customers, cutoff, horizon_days=60):
"""1 if the customer placed no order in the horizon after the cutoff."""
window = orders[(orders.order_date > cutoff) &
(orders.order_date <= cutoff + pd.Timedelta(days=horizon_days))]
return pd.Series(~customers.isin(window.customer_id), index=customers, name="churned").astype(int)Step 4: Train on the past, test on the more recent past
Instead of a random split, build the training set at one cutoff and the test set at a later one. That's exactly how the model will be used: trained on history, then applied to the next period. Make sure the training label window ends before the test cutoff.
CodeTwo cutoffs, no overlap
train_cut = pd.Timestamp("2026-03-31") # labels use April–May
test_cut = pd.Timestamp("2026-06-30") # labels use July–August
X_train = features(orders, train_cut)
y_train = labels(orders, X_train.index, train_cut)
X_test = features(orders, test_cut)
y_test = labels(orders, X_test.index, test_cut)
print(f"Churn rate: {y_train.mean():.0%} train, {y_test.mean():.0%} test")Step 5: Measure a simple baseline
Before training anything, check how well a single obvious rule ranks customers: "the longer since their last order, the more likely they've churned". Any model has to beat this baseline to be worth it.
CodeRecency as the baseline
from sklearn.metrics import roc_auc_score
print("Baseline ROC-AUC (recency only):", round(roc_auc_score(y_test, X_test.recency_days), 3))Step 6: Train a gradient-boosted model
Gradient-boosted trees such as LightGBM are the go-to choice for tabular data: accurate, fast on a normal CPU, and able to handle features on very different scales without preparation. Then check the number that matters for the retention team.
CodeTrain LightGBM and check the top 10%
from lightgbm import LGBMClassifier
model = LGBMClassifier(n_estimators=300, learning_rate=0.05, num_leaves=15, min_child_samples=50, verbose=-1)
model.fit(X_train, y_train)
risk = model.predict_proba(X_test)[:, 1]
print("Model ROC-AUC:", round(roc_auc_score(y_test, risk), 3))
top = pd.Series(risk, index=X_test.index).nlargest(len(risk) // 10).index
print(f"Churn rate in the riskiest 10%: {y_test[top].mean():.0%} (overall: {y_test.mean():.0%})")If the riskiest 10% churn at, say, three times the overall rate, the team's outreach is three times better targeted than contacting customers at random. That's the result to share with the business, not just the ROC-AUC.
If the model doesn't beat the recency baseline, that's a useful result too: use the simple rule for now, and add richer features such as support tickets, returns, discounts used or website visits, which carry signals that order history alone doesn't.
Step 7: Explain what drives the risk
People act on predictions they understand. SHAP shows how much each feature pushed each customer's risk up or down.
CodeFeature effects with SHAP
import shap
explanation = shap.TreeExplainer(model)(X_test)
shap.plots.beeswarm(explanation) # one dot per customer, per featureExpect recency and recent order counts to dominate. If a feature you didn't expect is at the top, check it for leakage before celebrating.
Step 8: Score customers every week
Before going live, retrain on the most recent complete window (cutoff = 60 days ago), so the model learns from the latest behaviour. Then, each week, compute features with today as the cutoff, with no labels needed, and hand the list to the team.
CodeWeekly scoring
today = pd.Timestamp.today().normalize()
X_now = features(orders, today)
scores = pd.Series(model.predict_proba(X_now)[:, 1], index=X_now.index, name="churn_risk")
scores.nlargest(len(scores) // 10).to_csv("at_risk_customers.csv")A scheduled job that runs this script weekly is all the infrastructure you need. There's no need for a real-time API. Every cloud has a cheap batch or scheduled-job service; see the cloud comparison.
Step 9: Prove it works with a control group
A good model doesn't automatically mean fewer customers leave. Only the retention action can do that. To measure it, randomly keep part of the high-risk list (for example 20%) out of the campaign. Compare how many churn in each group after 60 days. This A/B test is the only honest way to show the project's value.
Common pitfalls
- Building features with data from after the cutoff, for example "total orders this year" computed today.
- Including customers who left long ago. They inflate the churn rate and teach the model nothing useful.
- Random train/test splits across time, which let the model peek at the future.
- Measuring the model but never the campaign. Without a control group you can't tell if outreach helped.