Vision (modèles multimodaux)

An AI model's ability to read an image the way it reads text: a screenshot, a document photo, a chart, a mockup. You send the image in the same call as your question, and the model describes, extracts, compares or reasons over it. It's what moves an assistant from 'understands text' to 'understands what I show it'.

Strengths

Limitations

Best for

Official site

View on Coeurdar