Regression (Predict a Number)
Estimate a numeric value such as a price, a duration or a score.
Typical projects
Three ways to build it
Starter
AutoML regression (AutoGluon / Vertex AutoML)Lets you get a robust model without tuning anything by hand.
Best for: No or little data, or new to ML
AutoGluon · pandas
Standard
Gradient-boosted trees (LightGBM / XGBoost)Handles non-linear relationships and mixed feature types out of the box.
Best for: Some labeled data and Python experience
LightGBM · scikit-learn · pandas
Advanced
LightGBM with quantile objectives for prediction intervalsGives a range ('$310k–$345k'), which is often more useful than one number.
Best for: Lots of data and an experienced team
LightGBM · Optuna · MLflow
How success is measured
MAE (easy to explain: 'off by $X on average') and RMSE (punishes big misses)
The data you'll need
- Collect one row per item with the true value you want to predict.
- Look for outliers in the target — a few extreme values can dominate training.
- A few thousand rows is usually enough for a solid start.
Labeling
The target is usually a historical number (sold price, actual delivery minutes). Pull it from your transactional database.
Preparing the data
- Log-transform skewed targets like prices
- Impute missing values
- Encode categories
- Remove impossible values (negative prices, etc.)
Start with a baseline
Predict the mean or median of the target. Then try linear regression.
Evaluating the model
- Plot predicted vs actual
- Check error per segment (cheap vs expensive items)
- Report MAE in business units
Monitoring in production
- Input data drift
- Error once true values arrive
- Share of predictions outside the training range
Common pitfalls
- Extrapolation — models can't predict well outside the range they saw
- Using features only known after the fact
Example code
CodeQuick start
# pip install autogluon
from autogluon.tabular import TabularPredictor
import pandas as pd
df = pd.read_csv("houses.csv")
predictor = TabularPredictor(label="price", problem_type="regression").fit(df, time_limit=600)
print(predictor.leaderboard())CodeTrain your own model
# pip install lightgbm scikit-learn pandas numpy
import numpy as np, pandas as pd, lightgbm as lgb
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error
df = pd.read_csv("houses.csv")
X = pd.get_dummies(df.drop(columns=["price"]))
y = np.log1p(df["price"]) # log target for skewed prices
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42)
model = lgb.LGBMRegressor(n_estimators=1000, learning_rate=0.03)
model.fit(X_tr, y_tr, eval_set=[(X_te, y_te)])
pred = np.expm1(model.predict(X_te))
print("MAE:", mean_absolute_error(np.expm1(y_te), pred))
model.booster_.save_model("model.txt")