How the model ‘sees’ your image
This tool runs a vision-language model (BLIP/ViT-GPT2) that passes your image through a visual encoder to understand its content, then generates a natural-language sentence describing it — entirely in your browser. The output is a literal description of what’s visible, which is exactly what good alt text needs. It’s a draft, though, not a final answer: you’ll often want to add context the model can’t know (whose team, which product, what the chart proves).
Writing alt text the model gives you a head start on
The generated caption gets you 80% of the way to solid alt text. The remaining 20% is human judgment:
- Add context, trim detail. The model describes pixels; you know meaning. Keep what matters for the page, cut what doesn’t.
- Front-load the point. Lead with the most important element, since some users only hear the first few words before deciding to move on.
- Include text that appears in the image. If a graphic contains words (a sign, a chart label), put them in the alt text — the model may miss them.
Not every image needs a description
A subtle accessibility rule worth knowing: decorative images — borders, background flourishes, spacer graphics that add no information — should have empty alt text (alt=""), not a description. Describing purely decorative images just adds noise for screen-reader users. Reserve real alt text for images that carry meaning. The model will happily describe a decorative swoosh; your job is to decide whether that description helps anyone.
SEO and privacy
Descriptive alt text also helps Google Images understand and rank your visuals, so the same text serves accessibility and discoverability at once — just write it for humans first and the SEO benefit follows. And because the model (~100–200MB) runs locally after a one-time download, your images never upload, which matters when captioning product shots, client work, or anything you’d rather not send to a third-party API. The model is strongest on clear photographic subjects and weaker on abstract art or very cluttered scenes — review those captions extra carefully.