Category
Agents
Tool use, computer use and multi-step autonomy. Models built to act: calling tools, browsing, operating software and finishing multi-step jobs. Reach for an agent model when the job is a sequence rather than a single answer. Reliability under repeated matters more than brilliance on one turn, and stamina decides how many steps a run survives.
Coming soon
What to look for
What to look for in Agents
One model tracked today. These are the criteria that will separate the next ones.
- 01
Tool-calling reliability
One malformed tool call breaks a loop. Reliability under repeated calls matters more than single-shot brilliance.
- 02
Cost per completed task
Agents burn tokens across many steps. Evaluate cost per finished job, not per request.
- 03
Long-context stamina
Multi-step work accumulates context. Models that stay sharp at 100K+ tokens finish more tasks.
Explore nearby
Where to look next
Closest live categories today.