Datacurve's DeepSWE Shows Claude Opus 5 Beats Rivals at Half the Cost

Claude Opus 5 tops Datacurve's contamination-free DeepSWE benchmark at 74%, beating Fable 5 by 4 points at 45% lower cost per task

ByDatacurveDatacurve·
·
AuthorDatacurve
Read2 min
SubtopicLong Context
  • Claude Opus 5 tops DeepSWE v1.1 at 74%, the highest score on Datacurve's contamination-free long-horizon coding benchmark.
  • Opus 5 beats Fable 5 by 4 points (74% vs 70%) while costing 45% less per task ($11.84 vs $21.63 average).
  • Fable 5 narrates 65% of tool calls; Opus 5 only 11%, working silently and delivering a final summary instead.
  • GPT-5.6 Sol is a close second at 73%, completing tasks in only 61 steps vs Opus 5's 99, making it the most step-efficient top model.
  • DeepSWE uses harder tasks than SWE-bench Pro: 668 lines of code per solution vs 120, across 91 repos in 5 languages, with hand-written verifiers.
  • Claude models have a documented weakness: they miss parallel requirements (e.g., sync vs async branches) more than any other model family on DeepSWE.

Datacurve just added Claude Opus 5 to DeepSWE, its contamination-free long-horizon coding benchmark, and the results are striking: Opus 5 lands at 74% on v1.1, taking the top spot on the leaderboard. More interesting than the raw number is what it reveals about how Opus 5 behaves compared to its more expensive sibling, Claude Fable 5, and why this particular benchmark is worth paying attention to.

Why DeepSWE is a different kind of benchmark

Most public coding benchmarks are starting to saturate. On Scale AI's SWE-Bench Pro, the top models cluster within a narrow score band, making it nearly impossible to tell which agent will actually perform best in a real codebase. OpenAI's GPT-5 family, Anthropic's Claude Opus, and Google's Gemini Pro have clustered within a narrow band on Scale AI's SWE-Bench Pro leaderboard.

DeepSWE is a 113-task evaluation spanning 91 open-source repositories and five programming languages that produces a dramatically wider spread among the same frontier models. The key design choices that make this possible:

  • Contamination-free tasks: Tasks are written from scratch, not adapted from existing commits, making the benchmark contamination-free.
  • Harder prompts, larger solutions: DeepSWE prompts are roughly half the length of SWE-bench Pro's, yet the reference solutions average 668 lines of code added across 7 files, compared to SWE-bench Pro's 120 lines across 5 files.
  • Reliable verification: v1.1 updates execution and grading by scoring committed code in a clean, isolated environment. The verifier false positive rate is 0.3% versus SWE-Bench Pro's 8.5%, and the false negative rate is 1.1% versus 24.0%.

The leaderboard, in full

The DeepSWE v1.1 leaderboard has Claude Opus 5 at 74.0%, GPT-5.6 Sol at 72.7%, and 18 models ranked by long-horizon software engineering. Here is the full picture at max effort:

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves