Anomaly Detection
Spot unusual events: failing machines, suspicious logins, odd transactions.
Typical projects
Three ways to build it
Starter
Statistical rules + Isolation ForestUnsupervised, fast and easy to explain.
Best for: No or little data, or new to ML
scikit-learn · pandas
Standard
Isolation Forest / ECOD ensembles via PyOD with rolling featuresStrong unsupervised detectors with tunable alert volume.
Best for: Some labeled data and Python experience
PyOD · scikit-learn · pandas
Advanced
Autoencoder / sequence models + supervised model once labels accumulateCaptures complex multivariate patterns; supervised layer reduces false alarms.
Best for: Lots of data and an experienced team
PyTorch · PyOD · Kafka / Flink for streaming
How success is measured
Precision at a fixed alert budget (e.g. top 50 alerts/day) and recall on known incidents
The data you'll need
- Collect mostly 'normal' history — anomalies are rare by definition.
- Keep a list of known past incidents with timestamps; they become your test set.
- For sensors, keep raw high-frequency data plus aggregates.
Labeling
Usually unlabeled. Label a small set of confirmed incidents for evaluation; domain experts review alerts.
Preparing the data
- Normalise each signal
- Create rolling-window features (mean, std, rate of change)
- Separate by entity (per machine/user) since 'normal' differs
Start with a baseline
Simple statistical thresholds: alert when a value is > 3 standard deviations from its rolling mean.
Evaluating the model
- Check detections against known incidents
- Ask experts to review the top-N alerts
- Tune threshold to an alert volume the team can handle
Monitoring in production
- Alert volume per day
- Alert acknowledgement / false-positive rate
- Sensor dropouts
Common pitfalls
- Alert fatigue from too many false positives
- Treating seasonality (e.g. night vs day) as anomalies
Example code
CodeQuick start
import pandas as pd
df = pd.read_csv("sensor.csv", parse_dates=["ts"]).set_index("ts")
roll = df["temperature"].rolling("1h")
z = (df["temperature"] - roll.mean()) / roll.std()
alerts = df[z.abs() > 3]
print(alerts.head())CodeTrain your own model
# pip install scikit-learn pandas
import pandas as pd
from sklearn.ensemble import IsolationForest
df = pd.read_csv("sensor.csv", parse_dates=["ts"]).set_index("ts")
feats = pd.DataFrame({
"mean_1h": df["temperature"].rolling("1h").mean(),
"std_1h": df["temperature"].rolling("1h").std(),
"vibration": df["vibration"],
}).dropna()
model = IsolationForest(n_estimators=300, contamination=0.005, random_state=42)
feats["anomaly"] = model.fit_predict(feats) == -1
feats["score"] = -model.score_samples(feats.drop(columns="anomaly"))
print(feats.sort_values("score", ascending=False).head(20))