Anthropic's Claude Sonnet 5.5 Jumps 60 Points in Agentic Coding

Anthropic's mid-tier model gets a major overhaul with 30% faster generation, up to 30% lower cost per task, and a jump from 10.3% to 70.6% on Terminal-Bench.

·
·
  • Claude Sonnet 5.5 launches: 30%+ faster, up to 30% cheaper per task, same per-token pricing as Sonnet 5.
  • Terminal-Bench 4.0 jumps from 10.3% to 70.6%; near-Opus 5.5 performance on most benchmarks.
  • Pricing unchanged: $2/M input, $10/M output, $0.20/M cache reads on the Claude Platform.
  • Positioned for well-scoped work: bug fixes, slides, spreadsheets, UI polish, ticket handling.
  • First Sonnet with cyber safeguards, anti-distillation classifiers, and expanded preserved thinking.
  • Available now on AWS, Google Cloud, Azure via claude-sonnet-5-5; Haiku 5.5 coming soon.

Anthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family after Opus 5.5. It keeps Sonnet 5’s per-token pricing while generating output faster and using fewer tokens and tool calls to complete many tasks. Anthropic’s evaluations show the largest gains in agentic coding, computer use, and chart analysis.

A 60-point jump in terminal coding

Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5’s 10.3%. The benchmark measures agentic coding, in which a model works through terminal-based tasks using tools and multiple steps. Sonnet 5.5 also exceeds Opus 5.5’s reported 66.4% score on that evaluation.

On GDPval-AA, which covers practical work across several occupations, Sonnet 5.5 finishes two points behind Opus 5.5. It is also the first Sonnet model to complete Pokémon Red using screenshots as its only visual input, a test of sustained planning, visual interpretation, and action over a long sequence.

Benchmark results reported by Anthropic
Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5
Terminal-Bench 4.0 70.6% 10.3% 66.4%
CursorBench 4.0 55.5% 34.1% 57.8%
OSWorld 2.1, computer use 80.1% 57.0% 81.8%
Chartography, no tools 61.6% 15.6% 64.4%
Humanity’s Last Exam, with tools 64.5% 54.9% 67.7%

These launch figures place Sonnet 5.5 near Opus 5.5 on several evaluations and well ahead of Sonnet 5. CursorBench 4.0, which uses tasks drawn from real Cursor coding sessions, shows a 21.4-point gain over the previous Sonnet. Tool access and evaluation settings vary by benchmark, so comparisons are meaningful within each row rather than across different tests.

Fewer tokens cut task costs

Anthropic kept Sonnet’s API rates at $2 per million input tokens, $10 per million output tokens, and $0.20 per million prompt-cache read tokens. The company reports task-level savings of up to 30% because Sonnet 5.5 often finishes with fewer tokens and tool calls. Output generation is also more than 30% faster than Sonnet 5.

Claude Sonnet 5.5 API pricing
Usage Price per million tokens
Input $2.00
Output $10.00
Prompt-cache reads $0.20

Early integration partners reported lower token use and shorter completion times in their own evaluations:

  • Slack measured about 14% fewer output tokens on Slackbot evaluations without changing its prompts.
  • Balyasny Asset Management recorded approximately 121,000 tokens per answer, down from 497,000 with Sonnet 5, across a 2,441-task finance benchmark.
  • Box measured 2.4 times faster responses and 12% fewer total tokens than with the previous model.
  • Lovable recorded one-third fewer tool calls and roughly half as many shell runs per completed task.

For token-bound agents, total task cost depends on how quickly the model reaches a correct result. Better tool-call batching, fewer retries, and earlier termination can therefore reduce spending even when per-token rates remain unchanged.

A clock-of-clocks application created by Claude Sonnet 5.5
Anthropic’s clock-of-clocks example built with Sonnet 5.5.

Opus keeps the hardest assignments

Anthropic positions Sonnet 5.5 for well-scoped work such as bug fixes, ticket handling, document production, slides, spreadsheets, and routine multi-file edits. Opus 5.5 remains the stronger choice for open-ended assignments that require sustained judgment, evolving plans, or decisions without a single clear path.

Base44 compared the models across 118 real application builds and found that Sonnet 5.5 produced apps scoring level with Opus 5. Sonnet required an average of 3.6 iterations per build, compared with 7.7 for Opus 5. The team’s suggested workflow assigns architecture decisions to Opus and implementation work to Sonnet.

Early testers also reported gains in interface and presentation work. Anthropic says Sonnet 5.5 follows slide templates closely, adds useful visual polish to interfaces, and produces drafts that need less editing. In one internal test, the model created a 10-slide operating review from company earnings materials that two experts judged ready to send.

Effort becomes a budget control

Sonnet 5.5 supports the Claude 5.5 family’s effort controls, which trade token use and latency for additional reasoning. Claude Code and the Claude apps default to Medium effort, while the Claude Platform defaults to High. Lower settings return answers faster and consume fewer tokens; higher settings allow more reasoning and verification.

On CursorBench, Low-effort Sonnet 5.5 exceeded the best Sonnet 5 result at roughly one-tenth of the cost. That result applies to the cited evaluation, but it shows why effort should become an explicit variable in production testing rather than a fixed account-wide choice.

Cyber capability triggers tighter controls

Anthropic rates Sonnet 5.5’s cybersecurity capabilities as comparable to Opus 5, making this the first Sonnet release to use the company’s higher-tier cyber safeguards and fallbacks. Requests classified as higher risk can visibly route to Sonnet 5. Anthropic says the controls target a narrow set of requests, leaving routine software development largely unaffected. Biology safeguards remain the same as in Sonnet 5.

Sonnet 5.5 also introduces anti-distillation classifiers designed to prevent extraction of its reasoning for use in replicating or training another model. Expanded preserved-thinking controls bind reasoning state to the account that created it. Teams that transfer conversations between accounts or change accounts during a Claude Code session should review the preserved-thinking docs.

Check between_tools before upgrading

Sonnet 5.5 is available through Anthropic’s platforms and through Amazon Web Services, Google Cloud, and Microsoft Azure. Claude Platform users can select it with the model identifier claude-sonnet-5-5. Anthropic expects Haiku 5.5 to complete the family in the coming weeks.

  • Update the model identifier to claude-sonnet-5-5.
  • If Sonnet currently runs with thinking disabled, switch to the new between_tools setting before migration. This keeps up-front thinking disabled.
  • Retest latency, token budgets, tool-call limits, and completion quality at each effort level used in production.
  • For cybersecurity products, verify how higher-risk requests behave when the safety fallback activates.

Anthropic documents the thinking-setting change in its migration guide. Teams using Opus for bounded tasks such as bug fixes, ticket triage, slide generation, or multi-file edits now have a lower-cost candidate to evaluate. Production tests should use representative prompts, tools, and retry policies because those factors determine whether the benchmark and partner-reported savings carry over to a specific workload.

Trending
  • No trending articles

Comments

avatar

Next Reads