Classification answers one question well
This tool predicts what kind of thing an image depicts, using a vision model (MobileNet or ViT) trained on ImageNet’s 1.2 million images across 1,000 categories. Upload a photo and it returns its five most likely categories with confidence percentages. It’s fast, runs locally, and excels at one specific job: naming the prominent subject of a clear photo.
Why you get five answers, not one
Showing the top 5 isn’t hedging — it’s honest. Real photos are ambiguous, lighting varies, and many categories are visually close (terrier breeds, mushroom species, similar tools). A single forced answer would often be almost right; the top-5 list surfaces the correct label even when it’s the model’s second or third guess, and the confidence spread tells you how sure it is. A 95% top guess is confident; five guesses all near 20% means the image is genuinely hard to classify.
Classification vs detection — pick the right tool
| If you need… | Use |
|---|---|
| One label for the whole image | This classifier |
| To find and locate multiple objects | Object Detection |
| A written description of the scene | Image Captioning |
A common mistake is reaching for classification when you actually need detection — if your photo has several distinct objects and you want all of them located, classification will just name the most dominant one.
Getting the best results
The model is strongest on a single, prominent, well-lit subject filling most of the frame. It struggles with cluttered scenes (too many competing objects), abstract or heavily edited images, and subjects far outside its 1,000 categories. Crop tightly to your subject for a cleaner prediction.
Everything runs in your browser after a one-time model download (~20–50MB), so images are classified on your device and never uploaded — fine for personal photos, datasets, and anything confidential.