There is a version of this platform that would rank all fifty-four models by a single number and let you sort the table. It would be the most clickable thing on the site. We didn't build it, and the reason is worth explaining because it shapes how we present every benchmark.
A composite score hides the trade-off
A model that is excellent at reasoning is frequently only average at coding. Another is strong at multimodal tasks and unremarkable at everything text. Average those into one figure and you erase the single most important thing a benchmark can tell you: what this model is for.
A leaderboard answers "which model is best?": a question that has no answer without a task attached. The useful question is "best at what?", and a single score cannot carry it.
So we group scores by capability (reasoning, coding, multimodal) and show them separately. A model's shape matters more than its height.
Scores need a scale
A number floating on its own means nothing. Is 62 on a benchmark good? You cannot know without knowing what the field does. So every score sits beside a peer-median tick: the value has a scale, not just a magnitude.
- A model with no score for a benchmark shows an em dash, never a zero. Absence is not failure.
- Media and retrieval models are not scored on text benchmarks at all, and say so rather than showing blanks.
- Every figure carries the same caption, because these numbers are illustrative in this build.
Benchmarks are directional, not decisive
The most important thing to say about any benchmark is what it cannot do. Scores tell you which models are plausible for a task, worth putting on the shortlist. They do not tell you which one is right. Only your own evaluation, on your own data, at your real prompt depth, tells you that.
This is why the catalogue is built to get you to a shortlist quickly and then hand you off to the model detail pages, rather than to a ranking. The ranking would feel like an answer. The shortlist is an honest one.