Hello Model
← Model library

Object Detection

Find and locate objects in images or video with bounding boxes; count and track them.

Typical projects

Count people entering a store from CCTVDetect helmets on construction workersFind damaged areas on cars

Three ways to build it

Starter

Pretrained YOLO / open-vocabulary detector (YOLO-World, Grounding DINO)

Detect common objects or describe new ones in words — no training needed.

Best for: No or little data, or new to ML

Ultralytics · supervision

Standard

Fine-tuned YOLO (Ultralytics) on your labeled images

Excellent speed/accuracy, simple training CLI, exports to ONNX/TensorRT/CoreML.

Best for: Some labeled data and Python experience

Ultralytics · PyTorch · supervision

Advanced

RT-DETR / YOLO + tracking (ByteTrack) + TensorRT on edge GPUs

Production-grade real-time video pipelines with multi-camera tracking.

Best for: Lots of data and an experienced team

Ultralytics · NVIDIA DeepStream · TensorRT · ByteTrack

How success is measured

mAP (mean Average Precision) at IoU 0.5, plus FPS for video

The data you'll need

  • Collect frames covering different lighting, distances and occlusions.
  • Start with ~200–500 annotated images; more for small or rare objects.
  • Sample video frames sparsely so the dataset isn't full of near-duplicates.

Labeling

Draw boxes in CVAT, Label Studio or Roboflow. Use a pretrained model to pre-annotate and just correct it.

Preparing the data

  • Export labels in YOLO or COCO format
  • Augment with mosaic, scaling, brightness
  • Split by video/scene, not by frame

Start with a baseline

A COCO-pretrained YOLO model with no fine-tuning — it may already detect people, cars, etc.

Evaluating the model

  • mAP per class
  • Visualise predictions on held-out video
  • Measure FPS on the real target hardware

Monitoring in production

  • Detections per frame over time
  • Camera health (blank/blurred frames)
  • Periodic human review of sampled frames

Common pitfalls

  • Tiny objects need higher input resolution
  • Testing on frames from the same video clip as training

Example code

CodeQuick start
python
# pip install ultralytics
from ultralytics import YOLO
model = YOLO("yolo11n.pt")            # pretrained on 80 COCO classes
results = model("store_entrance.jpg")
people = [b for b in results[0].boxes if int(b.cls) == 0]
print("people detected:", len(people))
CodeTrain your own model
python
# pip install ultralytics
from ultralytics import YOLO

# data.yaml lists train/val image folders and class names
model = YOLO("yolo11s.pt")
model.train(data="data.yaml", epochs=100, imgsz=640, batch=16)
metrics = model.val()
print("mAP50:", metrics.box.map50)
model.export(format="onnx")   # or "engine" (TensorRT), "coreml", "tflite"

Ready to build one? Get a personalised plan →

Or read about LLM Assistant / RAG Chatbot next.