Claude Opus 5 Launches: Same Price, Bigger Benchmark Claims
24 Jul 2026
Anthropic has released Claude Opus 5, positioning it as the new default model on Claude Max and the strongest model available on Claude Pro. The launch arrives roughly two months after Opus 4.8 debuted on May 28, and follows a busy stretch for Anthropic's model lineup that saw Mythos 5, Fable 5, and Sonnet 5 all launch in June.
Sources differ on the exact release day: TechCrunch reports the launch happened on Friday, while The Verge states it was Thursday.
What's new in Opus 5
Anthropic is leaning heavily on benchmark comparisons to make its case. According to the report:
- On Frontier-Bench v0.1, Opus 5 surpasses all other models and more than doubles Opus 4.8's performance at a lower cost per task.
- On CursorBench 3.2 (max effort), Opus 5 lands within 0.5% of Fable 5's peak score at half the cost per task.
- On ARC-AGI 3, Opus 5 scores three times higher than the next-best model.
- On Zapier AutomationBench, its pass rate is roughly 1.5× the next-best model at the same cost.
- On OSWorld 2.0, it surpasses Fable 5's best result at just over a third of the cost.
- On FrontierCode 1.1, it approaches Fable-level performance at half the cost.
Domain-specific gains over Opus 4.8 include a 10.2 percentage-point improvement on organic chemistry tasks and 7.7 points on protein-related tasks. Anthropic also emphasized that Opus 5 is "much stronger at verifying its work and iterating carefully until it succeeds," and cited a case where the model wrote its own computer vision pipeline in response to an incomplete benchmark prompt.
Enterprise-facing results
Anthropic is marketing Opus 5 heavily toward enterprise customers, and third-party evaluations back some of the claims:
- Lovable's internal evaluations show Opus 5 up 22% over Opus 4.7 on the hardest agentic coding tasks.
- Box found Opus 5 outperforms Opus 4.8 by 8% overall, 11% in data analysis workflows, and 17% in due diligence workflows.
- On financial-modeling tasks, Opus 5 averaged 9 percentage points higher accuracy than Opus 4.8, using a third fewer turns/tool calls and 60% less time.
Pricing and safety changes
Despite the reported performance jump, Opus 5 keeps the same cost as Opus 4.8: $5 per million input tokens and $25 per million output tokens. A new Fast mode is launching in research preview at double the standard price. Notably, Opus 5 is slightly cheaper than OpenAI's recently released GPT-5.6.
On safety, Anthropic describes Opus 5 as its most aligned Opus model yet and the least susceptible to misuse. The company expects safety classifiers to engage 85% less often for Opus 5 than for Fable 5. Opus 5 is also not subject to the 30-day data retention policy that applies to Fable and Mythos.
However, the model still lags Mythos 5 on cybersecurity tasks, though the report doesn't specify the size of that gap. Opus 5's safeguards block it from scanning software binaries for vulnerabilities, but it is still permitted to search for vulnerabilities in source code — a distinction Anthropic frames as a stronger safeguard than Opus 4.8 had, but one that leaves some residual capability intact.
Anthropic is also rolling out Automatic Fallbacks, a system that reroutes requests to a less powerful model when a prompt trips the safety classifier. This mirrors broader industry dynamics: Fable 5 was previously pulled offline for a few weeks over government concerns before returning with stronger cyber safeguards, and OpenAI's GPT-5.6 underwent a roughly two-week trial limited solely to government-approved entities. Anthropic spokesperson Danielle Ghiglieri said the company "continue[s] to work with our government partners to conduct their own independent testing of our models. This includes Opus 5."
Why founders should care
- Startups running high-volume workloads may be able to access reported performance gains without paying more, since Opus 5's pricing matches Opus 4.8 — though actual savings will depend on task mix.
- Teams building coding agents, financial-modeling tools, or data-analysis products could plausibly see efficiency improvements based on the Lovable, Box, and financial-modeling figures cited, though these are vendor-reported and real-world variance is likely.
- Founders in security-sensitive domains should probably not assume parity with Mythos 5 on cybersecurity tasks and may want to evaluate alternative models for those specific workloads.
- Because Automatic Fallbacks can silently reroute prompts to a weaker model when safety classifiers trigger, founders relying on consistent output quality should test how this affects their specific use cases before going all-in.
- Anthropic's enterprise-first marketing (Box, Lovable case studies) may suggest pricing, support, and roadmap priorities skew toward larger organizations, which smaller teams should factor into vendor-selection decisions.
What's still unclear
The report notes several open questions: Anthropic hasn't detailed Frontier-Bench v0.1's methodology, the magnitude of the Mythos 5 cybersecurity gap, or specifics on context window size, multimodal capabilities, or architecture changes. TechCrunch's claim that Opus 5 "outperforms Fable 5" doesn't name the specific benchmarks involved, and there's no comparative detail on how government partners' independent testing results differ across models. The exact reasons behind Fable 5's earlier temporary removal also remain unexplained.
For now, Opus 5 is live and positioned as Anthropic's flagship — but founders evaluating it for production use should treat the benchmark claims as a starting point for their own testing, not a guarantee.