Object Detection
Find and locate objects in images or video with bounding boxes; count and track them.
Typical projects
Three ways to build it
Starter
Pretrained YOLO / open-vocabulary detector (YOLO-World, Grounding DINO)Detect common objects or describe new ones in words — no training needed.
Best for: No or little data, or new to ML
Ultralytics · supervision
Standard
Fine-tuned YOLO (Ultralytics) on your labeled imagesExcellent speed/accuracy, simple training CLI, exports to ONNX/TensorRT/CoreML.
Best for: Some labeled data and Python experience
Ultralytics · PyTorch · supervision
Advanced
RT-DETR / YOLO + tracking (ByteTrack) + TensorRT on edge GPUsProduction-grade real-time video pipelines with multi-camera tracking.
Best for: Lots of data and an experienced team
Ultralytics · NVIDIA DeepStream · TensorRT · ByteTrack
How success is measured
mAP (mean Average Precision) at IoU 0.5, plus FPS for video
The data you'll need
- Collect frames covering different lighting, distances and occlusions.
- Start with ~200–500 annotated images; more for small or rare objects.
- Sample video frames sparsely so the dataset isn't full of near-duplicates.
Labeling
Draw boxes in CVAT, Label Studio or Roboflow. Use a pretrained model to pre-annotate and just correct it.
Preparing the data
- Export labels in YOLO or COCO format
- Augment with mosaic, scaling, brightness
- Split by video/scene, not by frame
Start with a baseline
A COCO-pretrained YOLO model with no fine-tuning — it may already detect people, cars, etc.
Evaluating the model
- mAP per class
- Visualise predictions on held-out video
- Measure FPS on the real target hardware
Monitoring in production
- Detections per frame over time
- Camera health (blank/blurred frames)
- Periodic human review of sampled frames
Common pitfalls
- Tiny objects need higher input resolution
- Testing on frames from the same video clip as training
Example code
CodeQuick start
# pip install ultralytics
from ultralytics import YOLO
model = YOLO("yolo11n.pt") # pretrained on 80 COCO classes
results = model("store_entrance.jpg")
people = [b for b in results[0].boxes if int(b.cls) == 0]
print("people detected:", len(people))CodeTrain your own model
# pip install ultralytics
from ultralytics import YOLO
# data.yaml lists train/val image folders and class names
model = YOLO("yolo11s.pt")
model.train(data="data.yaml", epochs=100, imgsz=640, batch=16)
metrics = model.val()
print("mAP50:", metrics.box.map50)
model.export(format="onnx") # or "engine" (TensorRT), "coreml", "tflite"