The three flavors of ‘computer vision on an image’
People often conflate three related but distinct tasks. Choosing the right one saves a lot of confusion:
| Task | Answers | Output | This tool |
|---|---|---|---|
| Classification | ”What is this?” | One label for the image | Image Classifier |
| Detection | ”What and where?” | Labeled boxes per object | ✅ this tool |
| Segmentation | ”Which pixels?” | Pixel mask per object | — |
Object detection is the right tool when position and count matter — counting people in a crowd, locating products on a shelf, flagging whether a photo contains a particular object.
What it can find: the COCO 91
The DETR model is trained on the COCO dataset, which defines 91 everyday categories: people, animals (dog, cat, horse, bird), vehicles (car, bicycle, bus), and common objects (chair, laptop, phone, bottle, cup, book, food items). That coverage is broad enough for most real-world photos but is a hard boundary — anything outside those categories is invisible to the model. It won’t detect brands, specific people, text, or niche objects, because it was never taught them.
Reading the results
Each detected object comes with a bounding box, a label, and a confidence score. The score is your quality dial: high-confidence boxes are trustworthy, while low-confidence ones are the model hedging — useful to surface, but verify them. If you’re counting objects programmatically, set a confidence threshold appropriate to your tolerance for false positives vs. missed detections.
Practical uses and privacy
- Counting — people, vehicles, or items in a scene.
- Auto-tagging — generating object tags for a photo library.
- Pre-screening — checking whether an image contains a relevant object before deeper processing.
The model (~40–80MB) runs entirely in your browser after a one-time download, so your images stay on your device — important for surveillance frames, private photos, or proprietary product shots. Detection takes a few seconds per image depending on your hardware.