Cognition Brings GPT-6.1 Sol to Devin, Slashing Coding Costs by 57%

Cognition rolled GPT-6.1 Sol into Devin, matching the previous generation's score on its coding benchmark while cutting per-task costs by more than 80%.

·
·
Cognition Brings GPT-6.1 Sol to Devin, Slashing Coding Costs by 57%
  • GPT-6.1 Sol is live in Devin Desktop and Devin CLI.
  • Scores 60.4% on FrontierCode 1.1, matching GPT-6 Sol's 60.7%.
  • Costs $0.31 per task at medium effort, 81% cheaper than GPT-6 Sol at max.
  • Low-effort mode hits 58.1% for $0.21, up from 50.5%.
  • Highest score of any model under $0.30 per task on the leaderboard.
  • Runs 44% to 57% cheaper than GPT-6 Sol at every reasoning effort level.

GPT-6.1 Sol comes to Devin with lower task costs

Cognition has added OpenAI’s GPT-6.1 Sol to the Devin apps, including Devin Desktop and the Devin CLI. According to Cognition’s release notes, the model nearly matches its predecessors on coding quality while cutting the reported inference cost per benchmark task.

A flat score and a steep cost cut

On Cognition’s FrontierCode 1.1 Extended benchmark, GPT-6.1 Sol scores 60.4% at medium reasoning effort. GPT-6 Sol and GPT-5.6 Sol have reported headline scores of 60.7% and 60.6%, respectively. Reasoning effort controls how much computation the model can spend planning and checking its answer.

Selected FrontierCode figures published by Cognition
Model Reasoning effort Score Cost per task
GPT-6.1 Sol Medium 60.4% $0.31
GPT-6.1 Sol Low 58.1% $0.21
GPT-6 Sol Maximum 60.7% $1.66
GPT-6 Sol Low 50.5% Unstated

Cognition reports that GPT-6.1 Sol costs 44% to 57% less than GPT-6 Sol at matching reasoning settings, with scores remaining within one percentage point. The separate comparison between GPT-6.1 Sol at medium effort and GPT-6 Sol at maximum effort produces the larger 81% reduction, from $1.66 to $0.31, but uses different settings.

At low effort, GPT-6.1 Sol improves on GPT-6 Sol by 7.6 percentage points. Cognition says its 58.1% result is currently the highest FrontierCode score from any model costing less than $0.30 per task.

What “mergeable” means here

FrontierCode uses engineering tasks created by open source maintainers. An ensemble of unit tests, rubric-based checks, and automated verifiers evaluates each patch, and a run passes when the resulting code meets the benchmark’s mergeability criteria.

The 1.1 revision assigns zero credit to runs flagged for consulting internet sources that contain solutions. Cognition also audited the criteria that can block a patch from passing. These measures aim to reward code that satisfies project requirements and reduce the effect of leaked or memorized answers.

Because Cognition develops Devin and publishes FrontierCode, the results remain first-party measurements awaiting independent replication. The benchmark also has separate Main and Extended tracks; GPT-6.1 Sol’s 60.4% result comes from Extended. Reliable comparisons require the same track, reasoning setting, agent harness, and cost assumptions.

Agent loops magnify inference costs

Autonomous coding sessions can consume tokens across planning, repository searches, tool calls, retries, test runs, and self-verification. A 44% to 57% reduction at comparable benchmark quality can lower the cost of queued maintenance work and give a model router more room to use a frontier model.

Production expense will still depend on task length, context size, retry frequency, tool usage, and Devin’s billing model. FrontierCode’s per-task figure measures benchmark inference under Cognition’s setup, so teams should validate savings against complete session costs.

Cognition has already emphasized cost-aware model routing through Devin Fusion. The company describes Fusion as a multi-model system that reached 68.8% on FrontierCode 1.1 while reducing costs by 60% against comparable frontier performance. It has also reported Devin price cuts of up to 70% across coding modes. GPT-6.1 Sol gives that routing system another lower-cost option.

Five points behind, cents per task

Cognition’s Extended leaderboard places Anthropic’s Opus 5.5 at 65.3%, giving it a 4.9-point lead over GPT-6.1 Sol at medium effort. Cognition says flagship Opus and Fable runs have generally exceeded $1 per task at their strongest settings, compared with $0.21 to $0.31 for the two published GPT-6.1 Sol configurations.

That price difference creates a clear tradeoff for agent workloads: higher-scoring models may suit difficult tasks where each percentage point affects completion rates, while GPT-6.1 Sol may fit routine work executed at greater volume. Repository mix, latency, and human review time can change the economic result.

A disciplined Devin rollout

GPT-6.1 Sol is available now as a model choice in Devin Desktop and the Devin CLI. Teams moving workloads from GPT-6 Sol can begin with low effort for routine tasks and test medium effort on work that requires more planning or verification.

  • Run both models on the same representative repository tasks.
  • Keep prompts, tool permissions, acceptance tests, and time limits constant.
  • Measure completion rate, total session cost, latency, retries, and tool calls.
  • Record reviewer edits, regressions, and follow-up work after each run.
  • Retain automated tests and human review for production changes.

The published results make low-effort GPT-6.1 Sol a credible candidate for routine Devin work, with medium effort offering a higher benchmark score at modest additional cost. A team-specific evaluation will show whether those benchmark gains survive contact with its repositories, tools, and review standards.

Trending
  • No trending articles

Comments

avatar

Next Reads