OpenAI's GPT-6.1 Sol Cuts Benchmark Costs 76% but Runs 3x Slower
OpenAI's new GPT-6.1 Sol climbs to #7 on the Vals Index, uses far fewer tokens, but runs 2-3x slower and refuses cyber tasks.
- GPT-6.1 Sol ranks #7 on the Vals Index at 61.15%, up three spots from GPT-6 Sol.
- Pricing holds at $2/$10 per million tokens, but cached input cost halved and input token usage dropped 40-67%.
- Vibe Code Bench cost fell from $26.36 to $6.23 per run; IOI dropped from $2.70 to $1.26.
- MysteryMechanism gained 16.2 points by choosing simpler laws with fewer fitted parameters (3.6 vs 4.6).
- Runs 2-3x slower on agentic tasks: ~75 min per Vibe Code task, ~43 min per Legal Agent task.
- CyberBench collapsed to 39.3% due to stricter refusals; blocked 101 of 262 SRE Bench reverse-engineering tasks.
GPT-6.1 Sol Lowers Benchmark Costs and Slows Agent Runs
OpenAI has released GPT-6.1 Sol, an upgrade that improved coding and scientific-reasoning scores while reducing token costs in Vals AI’s benchmark runs. Those gains came with two constraints developers need to evaluate: agentic tasks took two to three times longer, and stricter security refusals blocked exploit generation and some reverse-engineering work.
Vals AI reported fewer cumulative input tokens, more reasoning tokens, and shorter agent trajectories than it observed with GPT-6 Sol. The results place GPT-6.1 Sol at No. 7 on the Vals Index with 61.15% accuracy, three positions above its predecessor.
Same rates, smaller token bills
OpenAI kept list prices at $2 per million input tokens and $10 per million output tokens. GPT-6.1 Sol supports a 1 million-token context window, which determines how much material a request can include, and up to 128,000 output tokens.
| Metric | GPT-6.1 Sol |
|---|---|
| Input price | $2 per million tokens |
| Output price | $10 per million tokens |
| Context window | 1 million tokens |
| Maximum output | 128,000 tokens |
| Cached input price | Half the previous rate |
Per-task savings came from lower token consumption. Across the Vals suite, GPT-6.1 Sol used 40% to 67% fewer input tokens while consuming roughly 2.25 times as many reasoning tokens, the internal tokens allocated to working through a problem before producing an answer.
| Benchmark | GPT-6 Sol cost | GPT-6.1 Sol cost | Reduction |
|---|---|---|---|
| Vibe Code Bench, full run | $26.36 | $6.23 | 76% |
| IOI, full run | $2.70 | $1.26 | 53% |
CodeMigration showed a similar price-to-performance gain. GPT-6.1 Sol improved by eight percentage points to 65.12% accuracy, ranking fourth at $6.51 per test. Claude Opus 5.5 ranked one position higher and cost $112.97 per test in the same evaluation.
Simpler laws lift scientific reasoning
MysteryMechanism produced the largest reported gain. The benchmark gives an agent a limited experiment budget, scaled to the problem’s dimension as 2d+1 trials, and asks it to infer a hidden mathematical law from the resulting data. GPT-6.1 Sol improved by 16.2 percentage points.
Both model versions used the available experiments and almost always submitted a valid law. GPT-6.1 Sol selected simpler structures, using a median of 3.6 parameters compared with 4.6 for GPT-6 Sol. Its predecessor more often chose flexible curves that fit the observations without capturing the underlying relationship.
In a task-by-task comparison, GPT-6.1 Sol solved 53 problems that GPT-6 Sol missed and failed on 17 that its predecessor solved, yielding a net gain of 36. On MysteryMechanism, the new model completed tasks in a median of 12 agent steps instead of 19 and used about 40% fewer tokens overall.
| Benchmark | Accuracy | Rank |
|---|---|---|
| IOI | 96.89% | No. 2 |
| ProofBench v1.1 | 99.00% | No. 4 |
| BioMysteryBench | 79.63% | No. 2 |
| Vibe Code Bench v1.1 | 88.93% | No. 6 |
| SRE Bench | 50.76% | No. 2 |
More reasoning raises latency
Higher reasoning use increased wall-clock time on long-running agent tasks. GPT-6.1 Sol ran two to three times slower than GPT-6 Sol, averaging about 75 minutes per Vibe Code Bench task and 43 minutes per Harvey Legal Agent task. That latency may limit interactive applications, synchronous developer tools, and CI jobs with strict time budgets.
Vals ran the model with reasoning_effort=max and default temperature, top-p, and top-k settings. Lower reasoning settings may reduce latency and accuracy, so teams should compare configurations against their own quality, cost, and response-time requirements.
Security filters block entire task classes
CyberBench v1.1 exposed a substantial change in refusal behavior. The benchmark asks a model to create a proof-of-concept exploit for a binary and then patch the vulnerability. GPT-6.1 Sol’s overall score fell to 39.29% from 78.0% for GPT-6 Sol.
| Security result | GPT-6 Sol | GPT-6.1 Sol |
|---|---|---|
| CyberBench v1.1 score | 78.0% | 39.29% |
| Exploit-generation refusal rate | 0% | 100% |
| Patching subtask | Unaffected | Unaffected |
The unchanged patching result suggests that the filter targeted offensive exploit generation in this evaluation while continuing to permit remediation. Similar refusals affected 101 of 262 SRE Bench tasks involving software reverse engineering, expanding the impact beyond explicit exploit development.
Choose by workload
GPT-6.1 Sol is suited to coding, mathematics, and scientific-reasoning workloads that can tolerate longer completion times. Its lower cumulative input use can reduce costs for multi-step agents, while the cheaper cached-input rate benefits repeated runs over the same codebase or document set.
- Coding and scientific research: Benchmark gains and lower per-task costs support migration testing.
- Interactive agents: Measure end-to-end latency at several reasoning settings before deployment.
- Security research: Test representative prompts because exploit development and reverse engineering may trigger refusals.
- Repeated codebase analysis: Evaluate cached-input savings alongside reasoning-token costs.
Developers can access the model through the standard OpenAI API. A production evaluation should track task accuracy, refusal rates, total token charges, and wall-clock latency because the benchmark results show meaningful movement in all four.