Hello Model
← Model library

Tabular Classification

Predict a category (yes/no, A/B/C) from rows of a spreadsheet or database.

Typical projects

Predict which customers will churnFlag fraudulent transactionsApprove or reject loan applications

Three ways to build it

Starter

AutoML (e.g. AutoGluon / Vertex AutoML / SageMaker Autopilot)

Tries dozens of models for you. Great when you are new or want a strong result fast.

Best for: No or little data, or new to ML

AutoGluon · pandas

Standard

Gradient-boosted trees (XGBoost / LightGBM)

The best accuracy-to-effort ratio for tabular data. Fast, CPU-only, explainable with SHAP.

Best for: Some labeled data and Python experience

LightGBM · scikit-learn · pandas · SHAP

Advanced

Tuned LightGBM/CatBoost ensemble + feature store

At scale, careful feature engineering and hyperparameter search beat fancier architectures.

Best for: Lots of data and an experienced team

LightGBM · CatBoost · Optuna · Feast

How success is measured

F1 score and ROC-AUC (use precision/recall if one kind of mistake is costlier)

The data you'll need

  • Export one row per entity (customer, transaction) with the outcome you want to predict as a column.
  • Make sure every feature is something you would actually know at prediction time — otherwise you get data leakage.
  • Aim for at least a few hundred examples of the rarest class.

Labeling

Labels usually already exist in your systems (e.g. 'cancelled_subscription = true'). Join them from your CRM / billing database.

Preparing the data

  • Handle missing values (impute median / 'unknown' category)
  • Encode categories (one-hot or target encoding)
  • Split by time if the data has dates, to mimic the future
  • Check class imbalance and consider class weights

Start with a baseline

Predict the majority class, then try a logistic regression. Any real model must beat both.

Evaluating the model

  • Confusion matrix on a held-out test set
  • Pick the decision threshold from business cost, not 0.5 by default
  • Explain predictions with SHAP to build trust

Monitoring in production

  • Feature data drift (distribution shift vs training)
  • Prediction rate per class
  • Real outcome vs prediction once labels arrive

Common pitfalls

  • Data leakage — a feature that secretly contains the answer
  • Optimising accuracy on imbalanced data (99% accuracy can mean useless)

Example code

CodeQuick start
python
# pip install autogluon
from autogluon.tabular import TabularPredictor
import pandas as pd

df = pd.read_csv("customers.csv")
predictor = TabularPredictor(label="churned").fit(df, time_limit=600)
print(predictor.leaderboard())
CodeTrain your own model
python
# pip install lightgbm scikit-learn pandas
import pandas as pd, lightgbm as lgb
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, roc_auc_score

df = pd.read_csv("customers.csv")
X = df.drop(columns=["churned"])
y = df["churned"]
X = pd.get_dummies(X)  # simple categorical encoding

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
model = lgb.LGBMClassifier(n_estimators=500, learning_rate=0.05, class_weight="balanced")
model.fit(X_tr, y_tr, eval_set=[(X_te, y_te)])

pred = model.predict_proba(X_te)[:, 1]
print("ROC-AUC:", roc_auc_score(y_te, pred))
print(classification_report(y_te, pred > 0.5))
model.booster_.save_model("model.txt")

Ready to build one? Get a personalised plan →

Or read about Regression (Predict a Number) next.