OpenAI's GPT-6.1 Sol Cuts Benchmark Costs 76% but Runs 3x Slower

OpenAI's new GPT-6.1 Sol climbs to #7 on the Vals Index, uses far fewer tokens, but runs 2-3x slower and refuses cyber tasks.

·
·
OpenAI's GPT-6.1 Sol Cuts Benchmark Costs 76% but Runs 3x Slower
  • GPT-6.1 Sol ranks #7 on the Vals Index at 61.15%, up three spots from GPT-6 Sol.
  • Pricing holds at $2/$10 per million tokens, but cached input cost halved and input token usage dropped 40-67%.
  • Vibe Code Bench cost fell from $26.36 to $6.23 per run; IOI dropped from $2.70 to $1.26.
  • MysteryMechanism gained 16.2 points by choosing simpler laws with fewer fitted parameters (3.6 vs 4.6).
  • Runs 2-3x slower on agentic tasks: ~75 min per Vibe Code task, ~43 min per Legal Agent task.
  • CyberBench collapsed to 39.3% due to stricter refusals; blocked 101 of 262 SRE Bench reverse-engineering tasks.

GPT-6.1 Sol Lowers Benchmark Costs and Slows Agent Runs

OpenAI has released GPT-6.1 Sol, an upgrade that improved coding and scientific-reasoning scores while reducing token costs in Vals AI’s benchmark runs. Those gains came with two constraints developers need to evaluate: agentic tasks took two to three times longer, and stricter security refusals blocked exploit generation and some reverse-engineering work.

Vals AI reported fewer cumulative input tokens, more reasoning tokens, and shorter agent trajectories than it observed with GPT-6 Sol. The results place GPT-6.1 Sol at No. 7 on the Vals Index with 61.15% accuracy, three positions above its predecessor.

Same rates, smaller token bills

OpenAI kept list prices at $2 per million input tokens and $10 per million output tokens. GPT-6.1 Sol supports a 1 million-token context window, which determines how much material a request can include, and up to 128,000 output tokens.

Metric GPT-6.1 Sol
Input price $2 per million tokens
Output price $10 per million tokens
Context window 1 million tokens
Maximum output 128,000 tokens
Cached input price Half the previous rate

Per-task savings came from lower token consumption. Across the Vals suite, GPT-6.1 Sol used 40% to 67% fewer input tokens while consuming roughly 2.25 times as many reasoning tokens, the internal tokens allocated to working through a problem before producing an answer.

Benchmark GPT-6 Sol cost GPT-6.1 Sol cost Reduction
Vibe Code Bench, full run $26.36 $6.23 76%
IOI, full run $2.70 $1.26 53%

CodeMigration showed a similar price-to-performance gain. GPT-6.1 Sol improved by eight percentage points to 65.12% accuracy, ranking fourth at $6.51 per test. Claude Opus 5.5 ranked one position higher and cost $112.97 per test in the same evaluation.

Simpler laws lift scientific reasoning

MysteryMechanism produced the largest reported gain. The benchmark gives an agent a limited experiment budget, scaled to the problem’s dimension as 2d+1 trials, and asks it to infer a hidden mathematical law from the resulting data. GPT-6.1 Sol improved by 16.2 percentage points.

Both model versions used the available experiments and almost always submitted a valid law. GPT-6.1 Sol selected simpler structures, using a median of 3.6 parameters compared with 4.6 for GPT-6 Sol. Its predecessor more often chose flexible curves that fit the observations without capturing the underlying relationship.

In a task-by-task comparison, GPT-6.1 Sol solved 53 problems that GPT-6 Sol missed and failed on 17 that its predecessor solved, yielding a net gain of 36. On MysteryMechanism, the new model completed tasks in a median of 12 agent steps instead of 19 and used about 40% fewer tokens overall.

Benchmark Accuracy Rank
IOI 96.89% No. 2
ProofBench v1.1 99.00% No. 4
BioMysteryBench 79.63% No. 2
Vibe Code Bench v1.1 88.93% No. 6
SRE Bench 50.76% No. 2

More reasoning raises latency

Higher reasoning use increased wall-clock time on long-running agent tasks. GPT-6.1 Sol ran two to three times slower than GPT-6 Sol, averaging about 75 minutes per Vibe Code Bench task and 43 minutes per Harvey Legal Agent task. That latency may limit interactive applications, synchronous developer tools, and CI jobs with strict time budgets.

Vals ran the model with reasoning_effort=max and default temperature, top-p, and top-k settings. Lower reasoning settings may reduce latency and accuracy, so teams should compare configurations against their own quality, cost, and response-time requirements.

Security filters block entire task classes

CyberBench v1.1 exposed a substantial change in refusal behavior. The benchmark asks a model to create a proof-of-concept exploit for a binary and then patch the vulnerability. GPT-6.1 Sol’s overall score fell to 39.29% from 78.0% for GPT-6 Sol.

Security result GPT-6 Sol GPT-6.1 Sol
CyberBench v1.1 score 78.0% 39.29%
Exploit-generation refusal rate 0% 100%
Patching subtask Unaffected Unaffected

The unchanged patching result suggests that the filter targeted offensive exploit generation in this evaluation while continuing to permit remediation. Similar refusals affected 101 of 262 SRE Bench tasks involving software reverse engineering, expanding the impact beyond explicit exploit development.

Choose by workload

GPT-6.1 Sol is suited to coding, mathematics, and scientific-reasoning workloads that can tolerate longer completion times. Its lower cumulative input use can reduce costs for multi-step agents, while the cheaper cached-input rate benefits repeated runs over the same codebase or document set.

  • Coding and scientific research: Benchmark gains and lower per-task costs support migration testing.
  • Interactive agents: Measure end-to-end latency at several reasoning settings before deployment.
  • Security research: Test representative prompts because exploit development and reverse engineering may trigger refusals.
  • Repeated codebase analysis: Evaluate cached-input savings alongside reasoning-token costs.

Developers can access the model through the standard OpenAI API. A production evaluation should track task accuracy, refusal rates, total token charges, and wall-clock latency because the benchmark results show meaningful movement in all four.

Trending
  • No trending articles

Comments

avatar

Next Reads