Skip to content

Category

Vision

Image and document understanding as input. Models that read what you show them: screenshots, documents, charts, photographs and interfaces. Reach for vision when the input is a picture and the output is words. If you need the opposite (), that is a different category. Resolution is the cost lever here: bigger images become more .

Coming soon

What to look for

What to look for in Vision

  1. 01

    Documents or scenes

    Document understanding (OCR, tables, forms) and natural-scene understanding are different skills. Test the one you need.

  2. 02

    Resolution costs tokens

    High-resolution image input can dominate your bill. Check how each model prices image tokens.

  3. 03

    Pair with a text model

    Many teams route vision requests to a multimodal model and everything else to a cheaper text model.