Hello Model
← Guides

Image classification · 5 min read

How to find defective products in photos

Spot scratched, cracked or stained products from a camera on the line, and send only the doubtful ones to a person.

What you’ll build

  • A model that scores every photo for defects, from a few hundred pictures
  • A threshold chosen from what a missed defect costs you
  • Pass, reject or “check by hand” for every part on the line

Step 1: Decide what counts as a defect, and what a miss costs

Write down which flaws make a part a reject, with a photo of each: a scratch longer than a few millimeters, any crack, a chipped edge, a stain. If two inspectors would disagree about a photo, the model will be confused by it too, so settle those cases first.

Then decide what matters more. A defective part that reaches a customer usually costs far more than a good part that gets a second look, so inspection models are tuned to catch nearly every defect (high recall) and accept some false alarms. Pick a target now, such as “catch 95% of defects”, and measure every model against it.

Step 2: Take photos the way the line will see them

Photograph parts with the camera, lighting and angle you’ll use in production: a model trained on bright studio photos will struggle under the factory lights. Fix the camera in place, light the parts evenly, and keep the background plain. Save the photos in two folders, photos/good/ and photos/defective/, and name each file after the part it shows, such as part0042_1.jpg.

A few hundred good parts and a few dozen defective ones are enough to start. No photos yet? This script draws realistic sample photos: 600 metal washers on a conveyor belt, each photographed twice under slightly different light, with about one in six scratched, cracked, chipped or stained.

CodeOptional: make sample photos
pythonRun in Colab
# pip install pillow numpy
import numpy as np
from pathlib import Path
from PIL import Image, ImageDraw, ImageFilter

rng = np.random.default_rng(7)
SIZE = 224

def make_part():
    """One metal washer: its size, and its defect (if any), at a fixed place on the part."""
    part = {"r_out": rng.uniform(70, 80), "r_in": rng.uniform(24, 30), "defect": None}
    if rng.random() < 0.17:                                    # about 1 part in 6 is defective
        part["defect"] = rng.choice(["scratch", "crack", "chip", "stain"])
        part["where"] = (rng.uniform(0, 2 * np.pi), rng.uniform(0.3, 0.8))   # angle, how far out
        part["shape"] = rng.uniform(0, 1, 8)
    return part

def photo(part):
    """A photo of the part on a conveyor belt; lighting, position and rotation change with every shot."""
    light = rng.uniform(0.75, 1.15)
    belt = (rng.normal(60, 6, (SIZE, SIZE)) * light).clip(0, 255).astype("uint8")
    img = Image.fromarray(np.stack([belt] * 3, -1))
    washer = Image.new("RGBA", (SIZE, SIZE))                  # drawn separately, then rotated onto the belt
    d = ImageDraw.Draw(washer)
    cx, cy = SIZE / 2 + rng.uniform(-14, 14), SIZE / 2 + rng.uniform(-14, 14)
    r_out, r_in = part["r_out"], part["r_in"]
    metal = int(rng.uniform(150, 185) * light)
    d.ellipse([cx - r_out, cy - r_out, cx + r_out, cy + r_out], fill=(metal, metal, metal + 6))
    d.ellipse([cx - r_in, cy - r_in, cx + r_in, cy + r_in], fill=(0, 0, 0, 0))
    for _ in range(60):                                        # harmless machining marks on every part
        a, r = rng.uniform(0, 2 * np.pi), rng.uniform(r_in + 3, r_out - 3)
        d.point((cx + r * np.cos(a), cy + r * np.sin(a)), fill=(metal - 30,) * 3)
    if part["defect"]:
        a, f = part["where"]; s = part["shape"]
        r = r_in + 8 + f * (r_out - r_in - 16)
        x, y = cx + r * np.cos(a), cy + r * np.sin(a)
        if part["defect"] == "scratch":
            b, n = s[0] * np.pi, 10 + 12 * s[1]
            d.line([x, y, x + n * np.cos(b), y + n * np.sin(b)], fill=(min(255, metal + 25),) * 3, width=1)
        elif part["defect"] == "crack":
            pts = [(x, y)]
            for k in range(6):
                x, y = x + 8 * (s[k] - 0.5), y + 8 * (s[(k + 3) % 8] - 0.5); pts.append((x, y))
            d.line(pts, fill=(int(metal * 0.55),) * 3, width=1)
        elif part["defect"] == "chip":
            x, y, c = cx + r_out * np.cos(a), cy + r_out * np.sin(a), 5 + 5 * s[0]
            d.ellipse([x - c, y - c, x + c, y + c], fill=(0, 0, 0, 0))
        else:   # stain
            c = 3 + 4 * s[0]
            d.ellipse([x - c, y - c * 0.7, x + c, y + c * 0.7], fill=(int(metal * 0.8), int(metal * 0.74), int(metal * 0.62)))
    washer = washer.rotate(rng.uniform(0, 360), center=(cx, cy))
    img.paste(washer, mask=washer)
    return img.filter(ImageFilter.GaussianBlur(rng.uniform(0.3, 1.0)))

# 600 parts, each photographed twice: photos/good/part0001_1.jpg, photos/defective/part0042_2.jpg, ...
for i in range(600):
    part = make_part()
    folder = Path("photos") / ("defective" if part["defect"] else "good")
    folder.mkdir(parents=True, exist_ok=True)
    for shot in (1, 2):
        photo(part).save(folder / f"part{i:04d}_{shot}.jpg", quality=90)
print({f.name: len(list(f.glob("*.jpg"))) for f in sorted(Path("photos").iterdir())})

Step 3: Keep both photos of a part on the same side

Split the photos into three groups: training (to learn from), validation (to make choices such as the threshold) and test (to measure the result once, at the end). If the same part appears in training and in test, the model only has to recognize it, and the test score is too good to be true. That’s data leakage, and it’s the most common mistake in image projects. Split by part, not by photo.

CodeSplit by part
pythonRun in Colab
# pip install scikit-learn pandas
import pandas as pd
from pathlib import Path
from sklearn.model_selection import GroupShuffleSplit

photos = pd.DataFrame({"path": [str(p) for p in sorted(Path("photos").glob("*/*.jpg"))]})
photos["defective"] = photos.path.str.contains("/defective/").astype(int)
photos["part"] = photos.path.str.extract(r"(part\d+)_")[0]     # both photos of a part share this

def split(df, test_size, seed=0):
    """Keep every photo of a part on the same side, so the test never shows a part the model has seen."""
    gss = GroupShuffleSplit(n_splits=1, test_size=test_size, random_state=seed)
    a, b = next(gss.split(df, groups=df.part))
    return df.iloc[a], df.iloc[b]

rest, test = split(photos, 0.25)
train, val = split(rest, 0.2)
for name, df in [("train", train), ("validation", val), ("test", test)]:
    print(f"{name:>10}: {len(df)} photos, {df.defective.mean():.0%} defective")

Step 4: Start with a pretrained model and a simple classifier

You don’t need to train an image model from scratch. A pretrained model such as EfficientNet has already learned to describe photos from millions of images. Use it to turn each photo into 1,280 numbers, then train a logistic regression on those numbers. This baseline trains in seconds and is surprisingly strong.

CodeThe baseline
pythonRun in Colab
# pip install tensorflow pillow
import numpy as np
import keras
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import recall_score

SIZE = 224
backbone = keras.applications.EfficientNetV2B0(include_top=False, pooling="avg", input_shape=(SIZE, SIZE, 3))

def load(paths):
    return np.stack([keras.utils.img_to_array(keras.utils.load_img(p, target_size=(SIZE, SIZE))) for p in paths])

def features(df):
    """1,280 numbers per photo that describe what's in it, from a model pretrained on millions of images."""
    return backbone.predict(load(df.path), batch_size=32, verbose=0)

X_train, X_val, X_test = features(train), features(val), features(test)
clf = LogisticRegression(max_iter=2000, class_weight="balanced").fit(X_train, train.defective)

def report(name, y, p_defect, threshold=0.5):
    flagged = p_defect >= threshold
    caught = recall_score(y, flagged)                     # share of defective parts we flag
    false_alarms = flagged[y == 0].mean()                 # share of good parts we flag by mistake
    print(f"{name}: catches {caught:.0%} of defects, flags {false_alarms:.1%} of good parts")

report("Baseline on validation", val.defective.values, clf.predict_proba(X_val)[:, 1])

On the sample photos, it catches 90% of defects and wrongly flags under 1% of good parts. It runs on an ordinary laptop: turning 1,000 photos into numbers takes about half a minute.

Step 5: Choose the threshold from the cost of a miss

The model gives each photo a probability of being defective. The threshold is where you draw the line, and 0.5 is rarely the right one. Lower it and you catch more defects but stop more good parts; raise it and the opposite happens. Choose it on the validation photos to meet your target from Step 1, then check it once on the test photos.

CodeA threshold that catches 95% of defects
pythonRun in Colab
def threshold_for(y, p_defect, catch=0.95):
    """The highest threshold that still catches the target share of defects."""
    return np.sort(p_defect[y == 1])[int(np.floor((1 - catch) * (y == 1).sum()))]

# Choose the threshold on the validation photos, then check it on the test photos.
threshold = threshold_for(val.defective.values, clf.predict_proba(X_val)[:, 1])
print(f"Threshold {threshold:.2f}")
report("Baseline on test", test.defective.values, clf.predict_proba(X_test)[:, 1], threshold)

On the test photos, the baseline now catches 93% of defects, but it flags 5.3% of good parts: one in twenty good parts would be stopped for nothing. Catching the last few defects is what costs the most false alarms.

Step 6: Learn what a good part looks like

Defects are small and rare, and new kinds keep turning up. So instead of learning what defects look like, learn what good parts look like, and flag anything unusual. This approach (often called PatchCore) looks at each photo in small patches, and gives the photo an anomaly score: how different its most unusual patch is from every patch seen on good parts. It needs no photos of defects at all, only of good parts.

CodeAn anomaly score from good photos only
pythonRun in Colab
# pip install scikit-learn
from sklearn.neighbors import NearestNeighbors

# Small details live in the middle layers of the network: one feature vector per 8×8-pixel patch.
mid = keras.Model(backbone.input, [backbone.get_layer("block3b_add").output, backbone.get_layer("block5e_add").output])

def patch_features(df):
    fine, coarse = mid.predict(load(df.path), batch_size=32, verbose=0)
    coarse = np.repeat(np.repeat(coarse, 2, axis=1), 2, axis=2)           # 14×14 → 28×28, to line up with fine
    both = np.concatenate([fine, coarse], axis=-1)
    both = keras.ops.convert_to_numpy(keras.layers.AveragePooling2D(3, strides=1, padding="same")(both))
    return both.reshape(len(df), -1, both.shape[-1])                       # photos × 784 patches × 160 numbers

# Memory of what good looks like: patches from good training photos only (a random sample keeps it fast).
rng = np.random.default_rng(0)
good = patch_features(train[train.defective == 0]).reshape(-1, 160)
memory = NearestNeighbors(n_neighbors=1).fit(good[rng.choice(len(good), 40_000, replace=False)])

def anomaly_score(df):
    """How unlike any good patch the most unusual patch of each photo is."""
    patches = patch_features(df)
    dist, _ = memory.kneighbors(patches.reshape(-1, patches.shape[-1]))
    return dist.reshape(len(patches), -1).max(axis=1)

s_val, s_test = anomaly_score(val), anomaly_score(test)
reject_from = threshold_for(val.defective.values, s_val)
report("Learned from good photos only, on test", test.defective.values, s_test, reject_from)

On the sample photos, this catches 94% of defects and flags only 2.0% of good parts, less than half the baseline’s false alarms, without seeing a single defect. That’s why it’s the usual starting point for inspection on a production line. Keep the defect photos you have: you need them to choose the threshold and to test the result.

Tip Fine-tuning the whole image model (transfer learning) can do even better once you have thousands of photos, including many defects, and a GPU. With a few hundred photos it’s easy to make things worse, so measure it against this step on the same test photos.

Step 7: Look at every mistake

Numbers tell you how often the model is wrong; photos tell you why. Put the missed defects and the false alarms side by side and look for what they share. On the sample photos, the misses are the faintest scratches and hairline cracks. Real projects often find something fixable: a reflection, a part placed off-center, a dirty lens.

CodeContact sheets of the mistakes
pythonRun in Colab
from PIL import Image

def contact_sheet(paths, file, size=160, columns=6):
    """Put photos side by side in one image, so you can see at a glance what the model gets wrong."""
    rows = max(1, -(-len(paths) // columns))
    sheet = Image.new("RGB", (columns * size, rows * size), "white")
    for i, p in enumerate(paths):
        sheet.paste(Image.open(p).resize((size, size)), ((i % columns) * size, (i // columns) * size))
    sheet.save(file)

flagged = s_test >= reject_from
missed = test[(test.defective == 1) & ~flagged]
false_alarms = test[(test.defective == 0) & flagged]
contact_sheet(missed.path, "missed.png")
contact_sheet(false_alarms.path, "false_alarms.png")
print(f"{len(missed)} missed defects in missed.png, {len(false_alarms)} false alarms in false_alarms.png")

Fix the cause, not the symptom: better lighting, a sharper camera or more photos of good parts in tricky positions usually help more than a different model.

Step 8: Pass, reject or check by hand

On the line, the model doesn’t have to decide everything on its own. Reject parts above the threshold, pass the ones with clearly normal scores, and send the doubtful middle to a person. People then see a small share of parts, and their decisions give you new labeled photos for the next version.

CodeDecide for each new photo
pythonRun in Colab
import joblib

# Save what the line needs: the memory of good patches and the two thresholds.
review_from = np.percentile(s_val[val.defective.values == 0], 90)     # the top 10% of scores for good parts
joblib.dump({"memory": memory, "reject_from": reject_from, "review_from": min(review_from, reject_from)}, "inspector.joblib")

def inspect(photo_path):
    """Pass, reject, or ask a person: the model only decides on its own when it's sure."""
    score = anomaly_score(pd.DataFrame({"path": [photo_path]}))[0]
    if score >= reject_from:
        return "reject", score
    if score >= review_from:
        return "check by hand", score
    return "pass", score

for path in [test[test.defective == 0].path.iloc[0], test[test.defective == 1].path.iloc[0]]:
    decision, score = inspect(path)
    print(f"{path}: {decision} (score {score:.1f})")
outcomes = pd.Series([inspect(p)[0] for p in test.path])
print(outcomes.value_counts(normalize=True).round(3).to_dict())

On the sample photos, three quarters of parts pass automatically, about one in five is rejected and 6% go to a person. The model can run on a small computer next to the camera (edge), or on a server; see the cloud comparison for hosting.

Once it’s running, keep an eye on it. Track the share of rejected parts each day, and check a sample of passed parts by hand every week. If the reject rate jumps, the parts may have changed, or the camera or lighting has (data drift). Retrain with fresh photos of good parts whenever something on the line changes.

Common pitfalls

  • Putting photos of the same part in training and in test, so the score is too good to be true.
  • Training on photos taken in different light, from a different angle or with a different camera than the line uses.
  • Keeping the default 0.5 threshold instead of choosing one from what a missed defect costs.
  • Measuring accuracy alone: with few defects, a model that passes everything still looks accurate.
  • Forgetting to retrain when the product, the camera or the lighting changes.

Want this tailored to your data, team and budget? Create a plan for your project

More on Image classification · Next guide: How to spot failing machines from sensor data