Glossary
Plain-English definitions of the terms used across the site.
- A/B test
- Showing two versions to random groups of users and comparing real outcomes.
- Accuracy
- Share of predictions that are correct. Misleading when one class is rare.
- Backtesting
- Testing a forecast by pretending to be at several past dates and comparing predictions to what actually happened.
- Baseline
- The simplest reasonable solution. Your model must beat it to be worth deploying.
- Chunks
- Small passages a document is split into so the most relevant pieces can be retrieved.
- Class imbalance
- When one category is much rarer than others (e.g. 1% fraud).
- Cold start
- The problem of recommending for brand-new users or items with no history.
- Data augmentation
- Creating extra training examples by slightly modifying existing ones (flip, crop, recolor).
- Data drift
- When live data starts looking different from the training data, which usually degrades accuracy.
- Data leakage
- When training data contains information that won't exist at prediction time — results look great in testing and fail in reality.
- Embeddings
- Lists of numbers that capture meaning; similar texts get similar embeddings.
- F1 score
- The balance between precision and recall in one number (0–1). Good for imbalanced classes.
- Fine-tuning
- Continuing to train a pretrained model on your own data so it specialises.
- Golden set
- A fixed list of test questions with known good answers used to measure quality after every change.
- Groundedness
- Whether the answer is actually supported by the retrieved documents, not made up.
- Hallucination
- When a language model states something confidently that is not true.
- Inference
- Using a trained model to make predictions.
- Linear regression
- Fits a straight line (or plane) through the data. The simplest baseline for predicting numbers.
- Logistic regression
- A simple, fast model that draws a straight boundary between classes. A great baseline.
- MAE
- Mean Absolute Error — on average, how far off predictions are, in real units.
- MAP
- Mean Average Precision — the standard accuracy score for object detection.
- MAPE
- Mean Absolute Percentage Error — average error as a percentage of the true value.
- MLOps
- Practices and tools for reliably deploying, monitoring and updating ML models.
- NDCG
- Ranking quality score that rewards putting the most relevant items at the top.
- Precision
- Of everything the model flagged, how much was actually correct.
- RAG
- Retrieval-Augmented Generation — look up relevant documents first, then let the LLM answer using them.
- Recall
- Of everything that should have been flagged, how much the model found.
- Recall@K
- Of the items a user actually liked, how many appear in the top K recommendations.
- RMSE
- Root Mean Squared Error — like MAE but punishes large errors more.
- ROC-AUC
- How well the model ranks positives above negatives across all thresholds. 0.5 = random, 1.0 = perfect.
- Seasonal naive
- A forecast that just repeats the value from the same point in the previous season (e.g. last week).
- TF-IDF
- Turns text into numbers by counting words, down-weighting very common ones.
- Transfer learning
- Starting from a model already trained on huge data and adapting it to your task with far less data.
- Vector database
- A database that finds items with the most similar embeddings quickly.
- WER
- Word Error Rate — share of words the transcription got wrong. Lower is better.
- Zero-shot
- Using a model on a task without giving it any training examples for that task.