Hello Model

Glossary

Plain-English definitions of the terms used across the site.

A/B test
Showing two versions to random groups of users and comparing real outcomes.
Accuracy
Share of predictions that are correct. Misleading when one class is rare.
Backtesting
Testing a forecast by pretending to be at several past dates and comparing predictions to what actually happened.
Baseline
The simplest reasonable solution. Your model must beat it to be worth deploying.
Chunks
Small passages a document is split into so the most relevant pieces can be retrieved.
Class imbalance
When one category is much rarer than others (e.g. 1% fraud).
Cold start
The problem of recommending for brand-new users or items with no history.
Data augmentation
Creating extra training examples by slightly modifying existing ones (flip, crop, recolor).
Data drift
When live data starts looking different from the training data, which usually degrades accuracy.
Data leakage
When training data contains information that won't exist at prediction time — results look great in testing and fail in reality.
Embeddings
Lists of numbers that capture meaning; similar texts get similar embeddings.
F1 score
The balance between precision and recall in one number (0–1). Good for imbalanced classes.
Fine-tuning
Continuing to train a pretrained model on your own data so it specialises.
Golden set
A fixed list of test questions with known good answers used to measure quality after every change.
Groundedness
Whether the answer is actually supported by the retrieved documents, not made up.
Hallucination
When a language model states something confidently that is not true.
Inference
Using a trained model to make predictions.
Linear regression
Fits a straight line (or plane) through the data. The simplest baseline for predicting numbers.
Logistic regression
A simple, fast model that draws a straight boundary between classes. A great baseline.
MAE
Mean Absolute Error — on average, how far off predictions are, in real units.
MAP
Mean Average Precision — the standard accuracy score for object detection.
MAPE
Mean Absolute Percentage Error — average error as a percentage of the true value.
MLOps
Practices and tools for reliably deploying, monitoring and updating ML models.
NDCG
Ranking quality score that rewards putting the most relevant items at the top.
Precision
Of everything the model flagged, how much was actually correct.
RAG
Retrieval-Augmented Generation — look up relevant documents first, then let the LLM answer using them.
Recall
Of everything that should have been flagged, how much the model found.
Recall@K
Of the items a user actually liked, how many appear in the top K recommendations.
RMSE
Root Mean Squared Error — like MAE but punishes large errors more.
ROC-AUC
How well the model ranks positives above negatives across all thresholds. 0.5 = random, 1.0 = perfect.
Seasonal naive
A forecast that just repeats the value from the same point in the previous season (e.g. last week).
TF-IDF
Turns text into numbers by counting words, down-weighting very common ones.
Transfer learning
Starting from a model already trained on huge data and adapting it to your task with far less data.
Vector database
A database that finds items with the most similar embeddings quickly.
WER
Word Error Rate — share of words the transcription got wrong. Lower is better.
Zero-shot
Using a model on a task without giving it any training examples for that task.