Cognition Adds Claude Sonnet 5.5 to Devin, Cutting Coding Costs by 30%

Cognition rolls Anthropic's newest mid-tier model into Devin Desktop and CLI, with a jump from 56.2% to 64.4% on FrontierCode at unchanged token pricing.

·
·
Cognition Adds Claude Sonnet 5.5 to Devin, Cutting Coding Costs by 30%
  • Cognition added Claude Sonnet 5.5 to Devin Desktop and CLI with no pricing change.
  • Scores 64.4% on FrontierCode 1.1 Main versus Sonnet 5's 56.2%, beating Fable 5.1 at extra-high effort.
  • Anthropic claims 30% faster output and up to 30% lower cost per task via fewer tokens and tool calls.
  • Lovable reports one-third fewer tool calls and roughly half as many shell executions on coding jobs.
  • Artificial Analysis found max-effort runs cost about 50% more per task than Sonnet 5.
  • Priced at $2/M input and $10/M output tokens, matching Sonnet 5 and half of Opus 5.5.

Cognition has added Claude Sonnet 5.5 to Devin Desktop and the Devin CLI, according to its release note. The model is live in both product selectors, requires no migration for existing Devin users, and uses Anthropic’s unchanged API rate card.

Long-running coding agents accumulate cost and latency across tokens, shell commands, and other tool calls. Fewer steps can lower wall-clock time and per-task cost under the existing token prices.

FrontierCode rises 8.2 points

FrontierCode 1.1 Main measures production-grade coding work through multistep tasks that require repository context and tools. Cognition reports the following results:

Model FrontierCode 1.1 Main
Claude Sonnet 5.5 64.4%
Claude Sonnet 5 56.2%

The gain is 8.2 percentage points, or about 14.6% relative to Sonnet 5. Sonnet 5.5 also edges past Fable 5.1 running at extra-high reasoning effort. Cognition maintains the benchmark, so the result provides a useful comparison within its harness rather than an independent assessment.

Fewer tool calls drive the savings

The reported rate card remains $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads. A cache read reuses previously supplied context, reducing the need to send the same prompt data again.

Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and can reduce total task cost by as much as 30%. The claimed savings come from fewer tokens and tool calls. Each avoided tool call removes an external action, such as a shell command or file operation, along with its network round trip and added context. These figures describe Anthropic API usage; Devin billing also depends on the customer’s plan.

Customer tests trace leaner agent loops

Customer evaluations cited for the release report lower tool use, token consumption, and execution time:

  • Lovable recorded about one-third fewer tool calls and roughly half as many shell executions across coding jobs.
  • Balyasny Asset Management measured about 121,000 tokens per answer, down from 497,000 with Sonnet 5, across a private suite of 2,441 finance tasks.
  • Box reported 2.4-times-faster results with 12% fewer total tokens.
  • Atlassian said customers could run Rovo Agents up to 30% faster than with Sonnet 5.
  • Zendesk reported 20% faster ticket processing and fewer incorrect decisions than the Claude models it currently uses in production.

The reports use different workloads, effort settings, and measurement methods, which prevents direct comparison across companies. Collectively, they indicate that reduced tool use accounts for much of the reported efficiency gain.

Scoped agent work is the target

Anthropic positions Sonnet 5.5 for well-defined daily tasks, including bug fixes and the production of documents, slides, spreadsheets, and design-sensitive output. On Terminal-Bench 4.0, an evaluation of agents working through terminal-based coding tasks, Anthropic reports a score of 70.6%, compared with 10.3% for the previous model. Sonnet 5.5 finished two points behind Opus 5.5 on GDPval-AA, which tests work drawn from real occupations.

Within Anthropic’s model family, Opus remains the higher-capability tier. Anthropic’s benchmarks place Sonnet 5.5 ahead of Opus 5.5 on agentic coding and attribute the result to its ability to spawn multiple agents within cost limits. That behavior maps directly to Devin jobs that distribute work across parallel sub-agents.

Maximum effort can raise the bill

Artificial Analysis recorded a more complicated cost profile. Its Intelligence Index, which combines ten evaluations, gave Sonnet 5.5 a score of 56 and Sonnet 5 a score of 38. The weighted cost per benchmark task reached $7.60 for Sonnet 5.5, compared with $5.09 for Sonnet 5.

The reasoning-effort setting controls how much computation and token generation the model can use before answering. At maximum effort, Sonnet 5.5 generated about 193,000 tokens per test task, the highest volume Artificial Analysis had measured. Each task cost $7.60 at that setting, about 50% more than a Sonnet 5 task. Savings near 30% therefore depend on the effort setting and on workloads where the model can reduce tokens and tool calls.

Completed tasks become the budget unit

Anthropic’s efficiency argument uses completed tasks as the relevant cost unit. Forecasts based only on tokens per request can miss gains from better planning, batched tool calls, and shorter execution paths. Devin’s long-horizon coding runs amplify those effects because every avoided shell invocation removes execution time and context from later prompts.

Teams can get the updated application from the Devin download page or install the CLI and select Sonnet 5.5. A controlled comparison should use representative repositories and the intended reasoning-effort setting while tracking task success, elapsed time, input and output tokens, cache reads, tool calls, and total cost.

Trending
  • No trending articles

Comments

avatar

Next Reads