Anthropic's Claude Opus 5 Triples ARC-AGI-3 Scores at Half the Frontier Price
Claude Opus 5 launches with a 3x ARC-AGI-3 lead, Fable 5-class intelligence at half the price, and Anthropic's strongest alignment scores yet

- Claude Opus 5 launches today on all paid plans and the Claude API, priced the same as Opus 4.8 at $5/$25 per million tokens.
- Near-Fable-5 intelligence at half the price — Opus 5 approaches Mythos-class performance while costing 50% less than Claude Fable 5.
- 3x ARC-AGI-3 lead over the next best model on the benchmark designed to test genuine fluid reasoning, not memorization.
- Most aligned Claude model to date — automated behavioral audit shows lowest rates of reckless or deceptive behavior across all Claude models.
- Fast mode ships with it, running at ~2.5x default speed — addressing a key complaint about Opus-class latency.
- Cybersecurity: capable but bounded — stronger than Opus 4.8 on security tasks, but substantially behind Mythos 5 at developing exploits.
Anthropic has released Claude Opus 5, the model the community had been tracking under the codename "Honeycomb," and the headline numbers are striking. It delivers intelligence close to the Mythos-class Claude Fable 5 at half the price, sets new state-of-the-art marks on coding and knowledge-work benchmarks, and triples the next best model's score on ARC-AGI-3. It ships today, priced identically to its predecessor Opus 4.8.
Where it fits in Anthropic's lineup
Anthropic's model stack runs from the budget Haiku tier through mid-tier Sonnet, up to Opus, with the premium Fable 5 above that. Mythos-class models sit above Opus in capability. Claude Fable 5 ($10/$50 per million tokens) is Anthropic's restored frontier tier, while Claude Mythos 5 is a trusted-access model with the same capability class but broader safeguards lifted for specific use cases.
Opus 5 slots below Fable 5 but approaches it in raw intelligence, while costing the same $5/$25 per million input/output tokens as Opus 4.8. That's a meaningful value proposition in a market where Claude Fable 5 already undercuts GPT-5.5 Pro ($30/$180) by 3–3.6x at the frontier tier.
The benchmark story
Anthropic claims Opus 5 is state-of-the-art on several coding and knowledge-work evaluations. The most striking number is on ARC-AGI-3, a benchmark worth understanding in depth.
ARC-AGI-3 is designed so that no amount of pretraining gives a meaningful advantage. The specific transformation rule in each puzzle is guaranteed to be novel, so memorization is useless. The benchmark was built by François Chollet and Mike Knoop's foundation, which set up an in-house game studio and created 135 original interactive environments from scratch. It's the closest thing the field has to a test of genuine fluid reasoning.
When ARC-AGI-3 first launched, Google's Gemini 3.1 Pro led at 0.37%, OpenAI's GPT-5.4 came in at 0.26%, Anthropic's Claude Opus 4.6 managed 0.25%, and humans solved 100% of environments. Against that baseline, Anthropic says Opus 5 scores three times as high as the next best model, a substantial relative jump on a benchmark specifically designed to resist gaming.
Gains on ARC-AGI-3 suggest models are developing capabilities that matter for real autonomous work: holding multiple abstract representations simultaneously, reasoning about transformation rules rather than just applying them, and generalizing from few examples.
Efficiency and alignment
Beyond raw benchmark scores, two other claims stand out. On efficiency, Anthropic says Opus 5 outperforms other models for a similar or lower cost per task, meaning it does more work per dollar than comparable frontier models, not just less than Fable 5.
On alignment, Anthropic ran an automated behavioral audit and says Opus 5 is their most aligned model to date, showing the lowest rates of reckless or deceptive behavior and the strongest adherence to Claude's Constitution. This continues a trend across the 4.x series. "Concerning behavior" scores measure a wide range of misaligned behavior, including cooperation with human misuse and undesirable actions the model takes on its own initiative, and Opus 5 posts the lowest figures yet on both.
Cybersecurity: capable but bounded
Opus 5 is stronger than Opus 4.8 on cybersecurity tasks, which is useful for developers doing vulnerability research and security audits. Anthropic is explicit, however, that it remains substantially behind Mythos 5 at developing exploits. Queries touching high-risk domains like cybersecurity, biology, chemistry, and model distillation are automatically detected and rerouted. The safeguards follow the same pattern Anthropic established with Fable 5's conservative safety classifiers: allow legitimate security work, block high-risk uses.
Availability and pricing
Opus 5 is available today across all paid plans and the Claude API:
- Claude Max: Default model
- Claude Pro: Strongest available model
- Claude API: $5/$25 per million input/output tokens, same as Opus 4.8
- Fast mode: Runs at approximately 2.5x the default speed
Fast mode is a meaningful addition. Speed has been a consistent complaint about Opus-class models, which trade latency for quality. A 2.5x speed option at the same price makes Opus 5 viable for interactive applications that previously had to settle for Sonnet.
Routing decisions in practice
For most teams, the model routing question now looks like this:
- Everyday coding and production workflows: Claude Sonnet 5 at $2–3/$10–15 per million tokens remains the price/performance default
- Complex reasoning, agentic coding, long-horizon tasks: Opus 5 at $5/$25, with near-Fable-5 intelligence
- Hardest frontier tasks, maximum autonomy: Fable 5 at $10/$50, where the benchmark premium justifies the cost
- Sensitive or restricted domains: Mythos 5 via Project Glasswing for approved partners
Opus 5 strengthens the middle tier between Sonnet 5 and Fable 5 considerably. Teams that have been routing to Fable 5 purely for its intelligence ceiling should re-evaluate: near-Mythos-class reasoning at half the API cost, with better alignment scores and a Fast mode option, changes the cost calculus for a wide range of production workloads.