How we evaluate models without a leaderboard
A leaderboard answers a question nobody actually has. We group benchmarks by capability and show them beside a peer median, because a model strong at reasoning is often mediocre at coding, and a composite score hides exactly the trade-off you're trying to see.