Tabular Classification
Predict a category (yes/no, A/B/C) from rows of a spreadsheet or database.
Typical projects
Three ways to build it
Starter
AutoML (e.g. AutoGluon / Vertex AutoML / SageMaker Autopilot)Tries dozens of models for you. Great when you are new or want a strong result fast.
Best for: No or little data, or new to ML
AutoGluon · pandas
Standard
Gradient-boosted trees (XGBoost / LightGBM)The best accuracy-to-effort ratio for tabular data. Fast, CPU-only, explainable with SHAP.
Best for: Some labeled data and Python experience
LightGBM · scikit-learn · pandas · SHAP
Advanced
Tuned LightGBM/CatBoost ensemble + feature storeAt scale, careful feature engineering and hyperparameter search beat fancier architectures.
Best for: Lots of data and an experienced team
LightGBM · CatBoost · Optuna · Feast
How success is measured
F1 score and ROC-AUC (use precision/recall if one kind of mistake is costlier)
The data you'll need
- Export one row per entity (customer, transaction) with the outcome you want to predict as a column.
- Make sure every feature is something you would actually know at prediction time — otherwise you get data leakage.
- Aim for at least a few hundred examples of the rarest class.
Labeling
Labels usually already exist in your systems (e.g. 'cancelled_subscription = true'). Join them from your CRM / billing database.
Preparing the data
- Handle missing values (impute median / 'unknown' category)
- Encode categories (one-hot or target encoding)
- Split by time if the data has dates, to mimic the future
- Check class imbalance and consider class weights
Start with a baseline
Predict the majority class, then try a logistic regression. Any real model must beat both.
Evaluating the model
- Confusion matrix on a held-out test set
- Pick the decision threshold from business cost, not 0.5 by default
- Explain predictions with SHAP to build trust
Monitoring in production
- Feature data drift (distribution shift vs training)
- Prediction rate per class
- Real outcome vs prediction once labels arrive
Common pitfalls
- Data leakage — a feature that secretly contains the answer
- Optimising accuracy on imbalanced data (99% accuracy can mean useless)
Example code
CodeQuick start
# pip install autogluon
from autogluon.tabular import TabularPredictor
import pandas as pd
df = pd.read_csv("customers.csv")
predictor = TabularPredictor(label="churned").fit(df, time_limit=600)
print(predictor.leaderboard())CodeTrain your own model
# pip install lightgbm scikit-learn pandas
import pandas as pd, lightgbm as lgb
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, roc_auc_score
df = pd.read_csv("customers.csv")
X = df.drop(columns=["churned"])
y = df["churned"]
X = pd.get_dummies(X) # simple categorical encoding
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
model = lgb.LGBMClassifier(n_estimators=500, learning_rate=0.05, class_weight="balanced")
model.fit(X_tr, y_tr, eval_set=[(X_te, y_te)])
pred = model.predict_proba(X_te)[:, 1]
print("ROC-AUC:", roc_auc_score(y_te, pred))
print(classification_report(y_te, pred > 0.5))
model.booster_.save_model("model.txt")