Step 1: Decide what to catch, and how many alerts people can check
Before any code, agree on two things with the people who will receive the alerts. First, what counts as catching a failure: here, an alert in the 3 days before a machine breaks down, so there’s time to plan a repair. Second, the alert budget: how many alerts the team can actually look at. Too many false alarms and people start ignoring all of them, which is called alert fatigue. This guide aims for at most one false alarm a day across the whole plant.
That gives you the scorecard for every approach: how many failures it caught, how many hours of warning it gave, and how many false alarms a day it raised.
Step 2: Collect sensor history and a list of past failures
You need two tables. One with a row per machine per hour: machine, time and each reading, such as temperature, vibration and pressure. And one with past failures from your maintenance log: machine and failed_at. You don’t need many failures to train on, because the model learns what normal looks like, but you need some to test against.
No data to hand? This script makes realistic sample data: 40 pumps over four months, working harder during two weekday shifts, with the odd glitchy reading. Half of them fail once, after 1.5 to 4 days of warning signs: either vibration creeps up (a worn bearing) or the pump runs hotter and loses pressure (a cooling problem).
CodeOptional: make a sample sensors.csv and incidents.csv
# pip install pandas numpy
import numpy as np
import pandas as pd
rng = np.random.default_rng(5)
hours = pd.date_range("2025-01-01", periods=120 * 24, freq="h")
n = len(hours)
# Two shifts on weekdays: pumps work harder from 6:00 to 22:00, and less at weekends.
load = np.where((hours.hour >= 6) & (hours.hour < 22), 1.0, 0.35) * np.where(hours.dayofweek < 5, 1.0, 0.6)
rows, incidents = [], []
for m in range(40):
machine = f"pump-{m + 1:02d}"
base_temp, base_vib, base_pres = rng.normal(55, 4), rng.normal(2.0, 0.3), rng.normal(6.0, 0.4)
def healthy():
return (base_temp + 12 * load + rng.normal(0, 0.8, n),
base_vib * (0.6 + 0.6 * load) + rng.normal(0, 0.08, n),
base_pres + 0.8 * load + rng.normal(0, 0.1, n))
temp, vib, pres = healthy()
glitch = rng.random(n) < 0.002 # sensor glitches: one wild reading, not a real problem
temp[glitch] += rng.choice([-25, 25], glitch.sum())
if m % 2 == 0: # half the pumps fail once in March or April
start, end = (62, 88) if m % 4 == 0 else (92, 118)
fail = int(rng.integers(start * 24, end * 24))
warn = int(rng.integers(36, 96)) # 1.5 to 4 days of warning signs
ramp = np.clip((np.arange(n) - (fail - warn)) / warn, 0, 1)
cause = rng.choice(["bearing", "cooling"])
if cause == "bearing":
vib += base_vib * 0.45 * ramp # vibration creeps up until the bearing seizes
else:
temp += 6 * ramp # the pump runs hotter and loses pressure
pres -= 0.5 * ramp
incidents.append((machine, hours[fail], cause))
down = slice(fail + 1, fail + 96) # no readings for 4 days while it's repaired
temp[down] = vib[down] = pres[down] = np.nan
after = slice(fail + 96, n) # then healthy again
t2, v2, p2 = healthy()
temp[after], vib[after], pres[after] = t2[after], v2[after], p2[after]
rows.append(pd.DataFrame({"machine": machine, "time": hours, "temperature": temp.round(1),
"vibration": vib.round(3), "pressure": pres.round(2)}))
pd.concat(rows).dropna().to_csv("sensors.csv", index=False)
pd.DataFrame(incidents, columns=["machine", "failed_at", "cause"]).to_csv("incidents.csv", index=False)
print(sum(len(r) for r in rows), "hourly readings from 40 pumps,", len(incidents), "known failures")Step 3: Split by time
Learn what normal looks like from the first two months, choose settings such as thresholds on the third, and test once on the fourth. Splitting by time, never at random, keeps the test honest: in real life you only ever know the past. On the sample data, the validation and test months have 10 failures each.
CodeLearn, choose, test
import numpy as np
import pandas as pd
sensors = pd.read_csv("sensors.csv", parse_dates=["time"]).sort_values(["machine", "time"])
incidents = pd.read_csv("incidents.csv", parse_dates=["failed_at"])
SIGNALS = ["temperature", "vibration", "pressure"]
# Learn what normal looks like from the first two months, choose settings on the third, test on the fourth.
train_end, val_end = pd.Timestamp("2025-03-01"), pd.Timestamp("2025-04-01")
in_val = lambda t: (t >= train_end) & (t < val_end)
in_test = lambda t: t >= val_end
print("Failures to catch: validation", in_val(incidents.failed_at).sum(), "· test", in_test(incidents.failed_at).sum())Step 4: Score alerts the way the team will see them
A real alert system doesn’t raise the same alarm every hour, so count alerts the way people will experience them: at most one per machine per day. An alert counts as useful if the machine fails within the next 3 days; every other alert is a false alarm.
CodeThe scorecard
WARNING = pd.Timedelta(hours=72) # an alert in the 3 days before a failure counts as catching it
def deduplicate(alerts, quiet=pd.Timedelta(hours=24)):
"""One alert per pump per day: after an alert, stay quiet about that pump for 24 hours."""
kept, last = [], {}
for i, a in alerts.sort_values("time").iterrows():
if a.machine not in last or a.time - last[a.machine] >= quiet:
kept.append(i)
last[a.machine] = a.time
return alerts.loc[kept]
def evaluate(name, alerts, period):
"""How many failures had an alert in the 3 days before them, how early, and how many alerts were false."""
alerts = deduplicate(alerts[period(alerts.time)])
failures = incidents[period(incidents.failed_at)]
warnings, useful = [], set()
for _, f in failures.iterrows():
before = alerts[(alerts.machine == f.machine) & (alerts.time <= f.failed_at) & (alerts.time > f.failed_at - WARNING)]
if len(before):
warnings.append((f.failed_at - before.time.min()) / pd.Timedelta(hours=1))
useful |= set(before.index)
days = sensors.time[period(sensors.time)].dt.normalize().nunique()
false_per_day = (len(alerts) - len(useful)) / days
warn = f", about {np.median(warnings):.0f} hours early" if warnings else ""
print(f"{name}: caught {len(warnings)} of {len(failures)} failures{warn}; {false_per_day:.1f} false alarms a day")
return len(warnings), false_per_dayStep 5: Start with a simple rule
The usual first rule is to alert when a reading is more than 3 standard deviations from the machine’s last 24 hours. It’s the baseline every smarter approach has to beat.
CodeThe baseline: a rolling 3-sigma rule
# Baseline: alert when a reading is more than 3 standard deviations from that pump's last 24 hours.
def rolling_z(df):
roll = df[SIGNALS].rolling(24, min_periods=12)
return (df[SIGNALS] - roll.mean()) / roll.std()
z = sensors.groupby("machine", group_keys=False).apply(rolling_z)
alerts = sensors.loc[(z.abs() > 3).any(axis=1), ["machine", "time"]]
evaluate("Rolling 3-sigma rule, validation", alerts, in_val);On the sample data it catches only 5 of the 10 failures and raises almost 3 false alarms a day. It has two problems. Each shift change looks like a jump, so it fires on perfectly healthy pumps. And because the window follows the readings, a slow rise in vibration over several days becomes the new normal before it ever looks unusual.
Step 6: Compare each machine with its own normal
What’s normal depends on the machine and the time: a warm pump at 14:00 on a Tuesday is fine, the same reading at 3:00 on a Sunday is not. So learn, from the training months only, each pump’s typical reading for every hour of a weekday and of a weekend, and how much its readings usually wander. Then measure each new reading against that, and smooth over 6 hours so one glitchy reading can’t raise an alarm.
CodeEach pump against its own normal
# What's normal depends on the pump, the hour and the day: a warm pump at 14:00 on a Tuesday is fine,
# the same reading at 3:00 on a Sunday is not. Learn each pump's normal from the training months only.
sensors["slot"] = sensors.time.dt.hour.astype(str) + np.where(sensors.time.dt.dayofweek < 5, "-weekday", "-weekend")
train = sensors[sensors.time < train_end]
normal = train.groupby(["machine", "slot"])[SIGNALS].median()
# How much each pump's readings usually wander around its normal: the yardstick for "unusual".
expected_train = normal.loc[list(zip(train.machine, train.slot))].to_numpy()
spread = (train[SIGNALS] - expected_train).groupby(train.machine).std()
def deviations(df):
"""How far each reading is from that pump's normal for that hour, in units of its usual spread,
smoothed over 6 hours so a single glitchy reading doesn't count."""
expected = normal.loc[list(zip(df.machine, df.slot))].to_numpy()
dev = (df[SIGNALS].to_numpy() - expected) / spread.loc[df.machine].to_numpy()
dev = pd.DataFrame(dev, columns=SIGNALS, index=df.index)
return dev.groupby(df.machine).transform(lambda s: s.rolling(6, min_periods=3).median())
dev = deviations(sensors)
alerts = sensors.loc[(dev.abs() > 3).any(axis=1), ["machine", "time"]]
evaluate("Each pump's own normal, validation", alerts, in_val);This catches all 10 failures in the validation month, about 35 hours before they happen, with fewer than one false alarm every two days. The model is no smarter than before; it just knows what normal looks like.
Step 7: Try a model, and test everything once
An Isolation Forest looks at all the signals together, plus how fast vibration is rising, and gives every hour an anomaly score. It needs no examples of failures: it learns from the training months, which are almost all normal. Choose its threshold on the validation month: the one that catches the most failures with at most one false alarm a day.
CodeAn Isolation Forest, with the threshold chosen on validation
# pip install scikit-learn
from sklearn.ensemble import IsolationForest
# Combine the signals, and add how fast vibration is rising: a pump that runs a little hotter AND loses
# a little pressure is more suspicious than either alone.
features = dev.assign(vibration_trend=dev.vibration.groupby(sensors.machine).diff(24)).dropna()
learn_from = features[sensors.loc[features.index, "time"] < train_end]
forest = IsolationForest(n_estimators=300, random_state=0).fit(learn_from)
score = pd.Series(-forest.score_samples(features), index=features.index) # higher = more unusual
def alerts_above(threshold):
return sensors.loc[score.index[score > threshold], ["machine", "time"]]
# Choose the threshold on the validation month: catch as many failures as possible
# with at most one false alarm a day, which is what the maintenance team said it can check.
results = {}
for share in [0.001, 0.002, 0.003, 0.005, 0.01]:
threshold = score.quantile(1 - share)
results[threshold] = evaluate(f"Top {share:.1%} of hours, validation", alerts_above(threshold), in_val)
threshold = max((t for t in results if results[t][1] <= 1), key=lambda t: (results[t][0], -results[t][1]))
print(f"Chosen threshold: {threshold:.3f}")Now run all three on the test month, once.
CodeThe test month
# The test month, once, with everything chosen beforehand.
evaluate("Rolling 3-sigma rule, test", sensors.loc[(z.abs() > 3).any(axis=1), ["machine", "time"]], in_test)
evaluate("Each pump's own normal, test", sensors.loc[(dev.abs() > 3).any(axis=1), ["machine", "time"]], in_test)
evaluate("Isolation Forest, test", alerts_above(threshold), in_test);On the sample data, the rolling rule catches 2 of 10 failures with 2 false alarms a day. Comparing each pump with its own normal catches all 10, about 38 hours early, with 0.3 false alarms a day. The Isolation Forest catches 9, with 0.4. The simple rule wins, and it’s easier to explain to the people who get the alerts, so that’s the one to run. That’s a common result: get the features right first, and only switch to a model when it clearly does better on your own data.
Step 8: Check every hour, and explain every alert
This is a small scheduled batch job: every hour, take the last few hours of readings, compare each machine with its normal, and send an alert that says which reading is off and by how much. “Pressure is 3.3 times its usual spread below normal” tells a technician where to look; a bare score doesn’t.
CodeThe hourly check
import joblib
# The simple rule won, so that's what runs. Save each pump's normal and its usual spread.
joblib.dump({"normal": normal, "spread": spread}, "pump_monitor.joblib")
def check(recent):
"""Every hour: compare each pump's latest readings with its normal, and explain any alert in plain words.
`recent` holds the last few hours for every pump (the 6-hour smoothing needs them)."""
recent = recent.assign(slot=recent.time.dt.hour.astype(str) + np.where(recent.time.dt.dayofweek < 5, "-weekday", "-weekend"))
d = deviations(recent)
latest = d.groupby(recent.machine).tail(1).dropna()
for i, row in latest.iterrows():
worst = row.abs().idxmax()
if abs(row[worst]) > 3:
direction = "above" if row[worst] > 0 else "below"
print(f"ALERT {recent.machine[i]} at {recent.time[i]:%b %d %H:%M}: {worst} is "
f"{abs(row[worst]):.1f} times its usual spread {direction} normal for this hour")
# Replay the hour the rule first raised the alarm before the first failure in the test month:
f = incidents[in_test(incidents.failed_at)].sort_values("failed_at").iloc[0]
rule = sensors.loc[(dev.abs() > 3).any(axis=1), ["machine", "time"]]
now = rule[(rule.machine == f.machine) & (rule.time > f.failed_at - WARNING)].time.min()
check(sensors[(sensors.time > now - pd.Timedelta(hours=12)) & (sensors.time <= now)])
print(f"{f.machine} failed on {f.failed_at:%b %d %H:%M} ({f.cause}), {(f.failed_at - now) / pd.Timedelta(hours=1):.0f} hours after this alert.")On the sample data, this warned about pump-27’s falling pressure 15 hours before it failed from a cooling problem.
Then keep it healthy. Track alerts a day, and ask technicians to mark each alert as real or false: those answers become your list of incidents for the next test. Relearn each machine’s normal every month or two, and right after any repair or change to the machine, or its new normal will look like a fault (data drift). Once you’ve collected a few dozen confirmed incidents, a supervised model trained on them can cut false alarms further.
Common pitfalls
- Comparing every machine with one plant-wide average, so normal differences between machines look like faults.
- Treating daily and weekly cycles (shift changes, weekends) as anomalies.
- Rolling windows that slowly accept a developing fault as the new normal.
- Choosing the threshold on the test month, so the result looks better than it will be.
- Raising more alerts than the team can check, until they stop reading them.