Cognition Slashes Devin Review Costs by 70% With Smarter Model Routing

Cognition slashes Devin pricing 30-70% across modes while claiming the top spot on its FrontierCode benchmark through smarter model routing and caching.

·
·
Cognition Slashes Devin Review Costs by 70% With Smarter Model Routing
Read4 min
TypeNews
  • Devin is now 30-40% cheaper in Fusion and Normal modes, 15-20% in Ultra, up to 70% in Devin Review.
  • Devin Fusion tops FrontierCode 1.1 Extended at 68.8, averaging $0.60 per task.
  • Model independence routes Opus 5.5, GPT-6 Sol, Astra, Luna, and SWE-2 to their strongest sub-tasks.
  • Harness now batches shell commands and runs independent tool calls in parallel to cut turns.
  • Prompt caching cuts processed-from-scratch tokens by ~71% in Cognition's five-turn example session.
  • SWE-2, post-trained from Kimi K3, powers cheaper runs but only inside Devin.

Cognition cuts Devin costs with model routing, batching, and caching

Cognition says its latest Devin update lowers average operating costs while preserving or improving the coding agent’s benchmark performance. The company reports reductions of 30% to 40% in Fusion and Normal modes, 15% to 20% in Ultra, and as much as 70% in Devin Review.

The savings come from three changes: routing each task to an appropriate model, batching tool calls, and keeping prompt prefixes stable to improve cache reuse. For development teams, those changes target the repeated inference costs that accumulate during long coding, testing, and debugging sessions.

Lower bills across every Devin mode

Mode Reported cost reduction Model strategy
Fusion 30% to 40% Pairs a lead model with a lower-cost model for supporting work
Normal 30% to 40% Selects models according to strengths in planning, testing, debugging, and migrations
Ultra 15% to 20% Uses higher-capability models selectively across the workflow
Devin Review Up to 70% Routes diff analysis, bug detection, and change categorization separately

Cognition also reports stronger benchmark results. Devin Fusion scored 68.8 on the Extended portion of FrontierCode 1.1 at an average cost of $0.60 per task, according to the company’s internal testing.

A model router sets the budget

Cognition calls its routing approach “model independence.” The harness, which coordinates the agent’s planning, tools, and context, can assign different parts of a coding job to different language models. Cognition uses Opus 5.5 and GPT-6 Sol for demanding reasoning, GPT-6 Astra for computer interaction, and GPT-6 Luna for lower-cost supporting tasks.

The available pool also includes Cognition’s SWE-2. The company says it post-trained the model with reinforcement learning using Kimi K3, Moonshot AI’s 2.8-trillion-parameter open model. SWE-2 scored 50.0% on FrontierCode 1.1 Main, one percentage point behind Fable 5.1 at a reported 64% lower cost. Cognition provides SWE-2 only through Devin, with no open weights or standalone API.

Two turns replace four

An agent harness determines how a model plans work, invokes tools, and incorporates their output. Newer models can request several actions at once, so Cognition changed Devin’s harness to batch related commands into a single tool call.

In Cognition’s example, Devin formats code, runs a linter, and executes tests together. Batching reduces the interaction from four model turns to two and sends 49% fewer tokens. The harness can also run independent commands in parallel, avoiding another model round trip between steps that do not depend on one another.

Stable prefixes cut repeat processing

Language model APIs generally process conversation history again on every turn because the models do not retain session state. Prompt caching lets providers reuse computation for an unchanged prefix, such as system instructions, repository context, and earlier messages. Any change within that prefix can prevent a cache hit.

Cognition says it redesigned Devin’s prompts to preserve shared prefixes and account for each provider’s cache lifetime and matching rules. In the company’s five-turn example, uncached processing consumes 93,000 tokens, while the cached version processes 27,000 from scratch, a 71% reduction.

Lower cache-read prices amplify those savings. Cognition reports that Opus 5.5 cache reads cost 60% less than Opus 5 reads, while GPT-6 Sol and Luna halve cache-read prices compared with their GPT-5.6 predecessors.

The benchmark boundary

FrontierCode is Cognition’s own benchmark, and the comparison results are vendor-reported and unaudited. The published figures therefore show Devin’s performance under Cognition’s test conditions, without establishing equivalent savings across other repositories, task types, or agent products.

SWE-2’s restricted availability creates another testing constraint. Developers cannot call it through a general API, inspect its weights, or isolate its performance from Devin’s routing and harness. The reported efficiency reflects the complete product rather than a model that teams can evaluate independently.

Put the savings through a repository trial

Teams evaluating Devin can test representative issues from their own repositories and measure completion rate, reviewer acceptance, wall-clock time, total tokens, cache-hit rate, and cost per accepted change. Those measurements reveal whether the reported inference savings persist across real codebases, long agent trajectories, and human review.

For developers building their own agents, Cognition’s implementation highlights three reusable controls: route work according to model strengths, batch independent tool calls, and preserve stable prompt prefixes. Each reduces repeated model processing without requiring a smaller task scope.

Trending
  • No trending articles

Comments

avatar

Next Reads