Anthropic's Claude Fable 5.1 Tops Every Frontier Model but Costs 20% More

Independent evaluations put Anthropic's new flagship at the top of the intelligence charts, but the win comes with a token-count tax that eats the cache savings.

·
·
Anthropic's Claude Fable 5.1 Tops Every Frontier Model but Costs 20% More
  • Claude Fable 5.1 tops the Artificial Analysis Intelligence Index at 66, ahead of Opus 5, GPT-5.6, and Grok 4.6.
  • Cache read pricing cut 75% from $1 to $0.25 per million tokens; base I/O pricing unchanged.
  • Despite the discount, Fable 5.1 costs 20% more per task than Fable 5 due to 1.7x output tokens.
  • New highs on HLE (59.1%), Terminal-Bench v2.1 (91.4%), SciCode (62.0%), and agentic work benchmarks.
  • Effectively tied with Opus 5 on GDPval-AA v2 and AA-Briefcase; wins analysis, loses presentation.
  • Available today on all Claude platforms; Mythos 5.1 twin restricted to trusted access.

Anthropic shipped Claude Fable 5.1 this week, and independent benchmarking from Artificial Analysis places it atop the leaderboard, ahead of every frontier model they track. The headline score comes with a twist: it costs more per task than the model it replaces, even after a 75% price cut on cached inputs.

A new number-one on the Intelligence Index

At maximum reasoning effort, Fable 5.1 scores 66 on the Artificial Analysis Intelligence Index, the highest they have measured. That puts it ahead of Claude Opus 5 at 63, the prior Fable 5 at 62, and both GPT-5.6 Sol and Grok 4.6 at 61. The jump over its predecessor is four points across the composite index.

The per-benchmark story is more nuanced. Fable 5.1 posts:

  • HLE: 59.1% (up from 55.5%)
  • Terminal-Bench v2.1: 91.4% (narrow lead)
  • SciCode: 62.0% (narrow lead)
  • Tau-3 Banking: +9 points over Fable 5
  • GDPval-AA v2: 1,853 Elo, +130 over Fable 5
  • AA-Briefcase: 1,694 Elo, +122 over Fable 5

Against Opus 5, the agentic knowledge-work leads on GDPval and AA-Briefcase are effectively tied once you factor in confidence intervals. Sub-scores show Fable 5.1 winning on analytical quality and rubric correctness but losing on presentation, worth knowing if your workflow involves customer-facing deliverables.

The pricing paradox

Anthropic kept base pricing identical to Fable 5, at $10 per million input tokens, $50 per million output tokens, and $12.50 per million cache write tokens. The big change is cache reads, which drop from $1 to $0.25 per million tokens, a 75% cut. VentureBeat notes this is aimed squarely at enterprise workloads where agentic loops re-read the same context repeatedly.

Yet Fable 5.1 at max effort costs $3.76 per Intelligence Index task, 20% more than Fable 5 at $3.14, and 1.6x Opus 5 at $2.34. The reason: it burns roughly 1.7x the output tokens. The cache discount saves about $1.40 per task, mostly on agentic evals where cache reads dominate input. Without the discount, Fable 5.1 would have cost around $5.16 per task.

If max effort is too expensive, the xhigh setting scores 65 at $2.72 per task, a $1.04 saving that gives up only one point on the index. Across its five effort settings, output token usage spans an 11x range, from 13.1M at low effort to 143.7M at max, with index scores from 58 to 66.

What actually shipped

The model has a 1 million token context window, accepts image and text inputs, and Anthropic lists a 128,000-token maximum output. Fable 5.1 is generally available on all Claude platforms today. The twin Mythos 5.1 model, which shares the same weights but ships with lighter safeguards, is restricted to trusted-access programs for cybersecurity and life sciences work.

One operational detail from the Artificial Analysis evaluation: they tested with Anthropic's default server-side fallback enabled, which reroutes safety-flagged requests to Opus 4.8 or Opus 5. That fallback served roughly 4% of output tokens across the index, so the headline scores include a small contribution from other models.

Higher accuracy, more confident wrong answers

On AA-Omniscience, Fable 5.1 attempts 93.4% of questions versus Opus 5's 87.8%, and records the highest accuracy Artificial Analysis has measured at 67.2%. It also hallucinates more when wrong: of questions it got incorrect, it still attempted an answer 72.6% of the time, versus 63.6% for Fable 5. The higher accuracy and higher hallucination rate cancel out, so its overall Omniscience Index score ends up level with the older model. If you care about calibrated refusals, that regression is worth testing on your own data.

When to reach for it

Fable 5.1 is positioned for coding and long-horizon agentic work. Anthropic describes it as built for jobs that take hours and span many applications: working through a backlog, operating a browser, or running unattended as a managed agent. Anthropic also claims Fable 5.1 fixes root causes rather than patching symptoms, and cites a case at investment firm Millennium where the model diagnosed a rare crash that had eluded engineers for years.

The practical read: if you already run agentic workloads heavy on cache reads, the 75% discount likely brings your bill down even though per-task costs on non-cached work go up. If you use Claude as a chat model for one-off requests, Fable 5.1 will cost more than Fable 5 for a modest quality bump, and Opus 5 remains the better cost-to-intelligence pick unless you specifically need the top of the leaderboard.

Comments

avatar