Skip to content

Category

Agents

Tool use, computer use and multi-step autonomy. Models built to act: calling tools, browsing, operating software and finishing multi-step jobs. Reach for an agent model when the job is a sequence rather than a single answer. Reliability under repeated matters more than brilliance on one turn, and stamina decides how many steps a run survives.

Coming soon

What to look for

What to look for in Agents

One model tracked today. These are the criteria that will separate the next ones.

  1. 01

    Tool-calling reliability

    One malformed tool call breaks a loop. Reliability under repeated calls matters more than single-shot brilliance.

  2. 02

    Cost per completed task

    Agents burn tokens across many steps. Evaluate cost per finished job, not per request.

  3. 03

    Long-context stamina

    Multi-step work accumulates context. Models that stay sharp at 100K+ tokens finish more tasks.