OpenAI's GPT-6.1 Sol Matches Flagship Performance at One-Fifth the Cost

OpenAI's mid-tier GPT-6.1 Sol matches its flagship Astra on coding and computer-use benchmarks while charging one-fifth the token price.

·
·
  • OpenAI released GPT-6.1 Sol, a mid-tier upgrade approaching GPT-6 Astra quality at one-fifth the price.
  • API pricing: $2 input, $10 output, and $0.10 cached input per million tokens.
  • Matches Astra on DeepSWE v1.1 coding tasks at roughly 20% of the cost per task.
  • Comes within 2.1 points of Astra on OSWorld 2.0 computer-use benchmark at one-seventh the cost.
  • Factual error rate drops from 11.4% to 7.7% at low reasoning effort versus GPT-6 Sol.
  • Available now in ChatGPT Work, Codex, and the API as gpt-6.1-sol; Ultrafast variant coming.

GPT-6.1 Sol Brings Near-Astra Performance to the Mid-Tier

OpenAI has released GPT-6.1 Sol, an upgrade to GPT-6 Sol for agentic coding, computer use, document analysis, and professional workflows. The company says it approaches GPT-6 Astra on several evaluations while charging substantially less per task.

Within OpenAI’s three-tier GPT-6 lineup, Astra remains the flagship, Sol occupies the mid-tier, and Luna provides the lowest-cost option. GPT-6.1 Sol keeps standard input at $2 per million tokens and output at $10 per million. Cached input falls to $0.10 per million tokens, 95% below standard input pricing and 50% below GPT-6 Sol’s cached rate.

GPT-6.1 Sol API pricing
Token type Price per million Pricing context
Standard input $2.00 Applies to uncached prompt tokens
Cached input $0.10 95% below standard input
Output $10.00 Applies to generated tokens
Comparison of GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna pricing tiers
OpenAI positions Sol between the premium Astra model and the lower-cost Luna tier.

Cached Context Changes the Math

Cached pricing applies to input tokens the API recognizes as reusable from earlier requests. At the listed rates, 10 million cached tokens cost $1, compared with $20 for the same volume of standard input. Actual savings depend on cache eligibility and hit rates because uncached input and generated output retain their standard prices.

Long-running agents often resend system instructions, repository context, tool definitions, or large documents across multiple calls. Those repeated prefixes make coding agents and document pipelines natural candidates for the lower cached rate. OpenAI’s repeated comparisons with Anthropic’s Opus 5.5 also place Sol in direct competition for document analysis and business automation workloads.

Near-Astra Scores at Lower Task Costs

OpenAI cites results from several external benchmarks to support its near-Astra positioning. Reasoning effort controls how much inference-time work the model performs, with higher settings generally consuming more tokens. The reported cost-per-task figures combine each run’s token usage with the applicable API rates, so production costs will vary by workload.

  • Agentic coding: On DeepSWE v1.1, which tests software-engineering work in real codebases, GPT-6.1 Sol matches GPT-6 Astra at about one-fifth of the cost per task. It also exceeds GPT-6 Sol’s best score by 6.4 percentage points at a lower reasoning setting and cost.
  • Document question answering: On GDP.pdf, which covers complex finance, healthcare, and legal documents, GPT-6.1 Sol scores above the tested Opus 5.5 configuration with fallbacks enabled. It costs less than half as much per task across the evaluated reasoning settings.
  • Business automation: On AutomationBench, GPT-6.1 Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort while costing about one-third as much.
  • Computer use: On the OSWorld 2.0 offline set, GPT-6.1 Sol comes within 2.1 percentage points of Astra at maximum reasoning effort and costs about one-seventh as much per task.
  • Scientific research: At maximum effort, GPT-6.1 Sol averages $5.47 per task, compared with $23.21 for Opus 5.5 and $23.80 for Astra. Astra retains the highest score on this evaluation at 68.1%.

On OpenAI’s adversarial factuality set, GPT-6.1 Sol shows its largest gain over GPT-6 Sol at low reasoning effort. The share of responses containing a factual error falls from 11.4% to 7.7%, a relative reduction of about 32%. The dataset uses conversations in which users flagged an earlier model’s mistake, so its absolute error rates are higher than those expected across typical traffic.

Tool Failures Get Clearer Handling

OpenAI also reports lower failure rates for disclosing unavailable tools, following explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. The broken-search evaluation checks whether the model tells the user that search failed. According to the system card addendum, GPT-6.1 Sol fails to disclose the problem in 2.1% of cases, compared with 4.9% for GPT-6 Sol and 1.5% for Astra.

Rollout, Model ID, and Migration Tests

OpenAI says the launch rollout covers paid ChatGPT plans, Codex, and the API. Availability differs by product surface:

Surface Availability Details
OpenAI API Available Use gpt-6.1-sol
ChatGPT Work and Codex Rolling out Plus, Pro, Business, Enterprise, and Edu plans
Chat Not yet available No release date provided
Codex Ultrafast Coming after launch Up to eight times faster token generation

A representative migration test can establish whether Sol’s lower price offsets any quality difference for a specific production workload. Teams moving from Astra can evaluate:

  1. Task completion rates against fixed acceptance criteria.
  2. Total token cost, including cache hit rates and generated output.
  3. Median and tail latency at the required reasoning settings.
  4. The share of tasks that still require an Astra fallback.
  5. Tool-use failures, permission violations, and unsupported claims.

The reported results make GPT-6.1 Sol a candidate default for coding agents, document pipelines, and desktop automation. Astra retains an advantage on the cited scientific-research evaluation and may remain appropriate where measured quality gains cover its higher cost. Cache-heavy workloads receive the largest direct savings when repeated prompt content consistently qualifies for cached pricing.

Trending
  • No trending articles

Comments

avatar

Next Reads