Claude Opus 5 Tops Artificial Analysis AI Index
28 Jul 2026
Claude Opus 5 (Adaptive Reasoning, Max Effort) currently sits atop the Artificial Analysis Intelligence Index, scoring 61 out of a possible range among 170 evaluated models—the highest mark on the leaderboard as of this report.
The numbers
- Claude Opus 5: Intelligence Index score of 61, ranked #1 among 170 models
- GLM-5.2 (max): score of 51, the top-ranked open weights model
- Open weights models: 94 of the 170 total evaluated
- Reasoning models: 126 evaluated
- Mercury 2: fastest model at 938.7 tokens/second
- Nova Micro: cheapest at $0.03 per 1M tokens
- Gemini 2.5 Flash-Lite (Non-reasoning): lowest time to first token at 0.35 seconds
- Standardized prompt evaluations were run across a broader set of 586 models
What this means
Claude Opus 5 leads by a meaningful margin—its score of 61 is 10 points ahead of GLM-5.2 (max), the best-performing open weights model at 51. That gap could be relevant for teams weighing proprietary versus open weights options for reasoning-heavy applications, though the report doesn't specify how large a real-world performance difference this represents.
Beyond raw intelligence scoring, the index highlights that no single model dominates every dimension. Mercury 2 leads on raw speed (938.7 tokens/second), Nova Micro leads on cost ($0.03 per 1M tokens), and Gemini 2.5 Flash-Lite leads on latency (0.35 seconds to first token). Founders building latency-sensitive or cost-sensitive products may find these models more relevant than Opus 5, depending on the tradeoff that matters most for their use case.
One notable inconsistency in the underlying data: the Intelligence Index covers 170 models, while a separate standardized prompt evaluation spans 586 models. The report does not clarify why these totals differ, so founders should confirm which specific benchmark set applies before drawing conclusions about a given model's relative standing.
Why founders should care
- If your product depends heavily on reasoning quality, Opus 5's top score suggests it's likely worth evaluating for reasoning-intensive tasks—though benchmark leadership doesn't guarantee it will outperform on your specific workload.
- If you're building on open weights infrastructure, GLM-5.2 (max) is probably the strongest current option in that category based on this index, but it still trails Opus 5 by 10 points.
- Teams optimizing for speed or cost may find Mercury 2 or Nova Micro more practically relevant than the top-ranked reasoning model, since intelligence rank and operational efficiency are measured separately.
- Leaderboard positions can shift quickly as new models are added, so founders should treat current standings as a snapshot rather than a durable signal.
What's missing
The report leaves several open questions: there's no methodology detail on how the Intelligence Index score is calculated, no timestamp on when Opus 5 was released or how long it's held the top spot, and no explanation of how the "Adaptive Reasoning, Max Effort" configuration compares to other Opus 5 settings. Founders should treat this ranking as a directional signal rather than a definitive verdict, and validate any model choice against their own task-specific benchmarks before committing.