OpenAI Brings 8x Faster Ultrafast Inference to GPT-6.1 Sol

OpenAI's premium speed tier now covers its mid-tier reasoning model, hitting up to 8x standard token generation speeds for latency-critical agentic workloads.

·
·
·
Read4 min
TypeNews
TopicLlms · Api
  • OpenAI rolled out Ultrafast mode for GPT-6.1 Sol in API, Codex, and ChatGPT Work.
  • Up to 8x faster token generation than Sol Standard, reaching around 300 tokens per second.
  • API pricing is $12 per million input tokens and $60 per million output tokens, 6x Standard.
  • In Codex and ChatGPT Work, access requires Pro 500, usage-based Enterprise, or credit-based Edu plans.
  • Supported in all regions with US and EU data residency; EU residency also added for Sol Fast and Luna Fast.
  • Built for outage debugging, agents navigating apps, and live experiences where latency matters.

OpenAI brings Ultrafast inference to GPT-6.1 Sol

OpenAI has extended its Ultrafast tier to GPT-6.1 Sol across the API, Codex, and ChatGPT Work. The service promises up to eight times faster token generation than Standard mode, targeting interactive agents, live developer tools, and incident-response workflows.

OpenAI introduced Ultrafast for GPT-6 Astra at its DevDay announcement in 2026 and previewed support for Sol. That support is now available, subject to account and plan eligibility.

An 8x lane for the same model

Ultrafast runs the same GPT-6.1 Sol model on faster inference infrastructure. API clients select it by setting service_tier to ultrafast in a Responses API request.

OpenAI documents generation speeds of up to 300 tokens per second for Astra in Ultrafast mode. For Sol, the company advertises up to an eightfold improvement over Standard mode without publishing a separate absolute maximum.

The premium is 6x per token

API calls to GPT-6.1 Sol cost six times more in Ultrafast mode than in Standard mode:

Service tier Input per 1M tokens Output per 1M tokens
Standard $2 $10
Ultrafast $12 $60

GPT-6 Astra Ultrafast costs $60 per million input tokens and $300 per million output tokens. Sol therefore carries one-fifth of Astra’s Ultrafast token price.

Plans, regions, and limits

Codex and ChatGPT Work provide Ultrafast access through the following plans:

  • Pro 500
  • Eligible usage-based Enterprise contracts
  • Credit-based Edu plans

Enterprise administrators must enable the tier before users can select it. OpenAI says Ultrafast is available in every supported region, including deployments with US or EU data residency. The same rollout added EU data residency for GPT-6.1 Sol Fast and GPT-6 Luna Fast.

Default API throughput limits increase with the customer’s usage tier:

Usage tier Default token limit
Build 500,000 tokens per minute
Launch 1,000,000 tokens per minute
Grow 5,000,000 tokens per minute

WebSockets protect the latency gain

OpenAI recommends WebSockets for agents that make frequent tool calls because a persistent connection reduces request overhead between turns. HTTP remains available through the SDK, though repeated request cycles can consume part of the latency saved during generation.

Because the advertised eightfold improvement measures token generation, teams should benchmark complete workflows. End-to-end latency also includes connection overhead, time to the first token, tool execution, and application processing.

Python clients can select GPT-6.1 Sol Ultrafast with a standard Responses API call:

python
from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    service_tier="ultrafast",
    input="Explain why the sky is blue in one sentence.",
)

print(response.output_text)

Multi-turn agents can maintain one WebSocket session and pass previous_response_id between turns to link responses without repeatedly sending the full conversation history.

Why Sol fits interactive work

OpenAI positions GPT-6.1 Sol close to GPT-6 Astra on agentic coding, computer use, and professional tasks while charging one-fifth of Astra’s token price. Sol supports a 1,050,000-token context window, including up to 922,000 input tokens and 128,000 output tokens.

OpenAI identifies three primary workload categories:

  • Incident response: Engineers can iterate faster while diagnosing production failures and testing fixes.
  • Computer-use agents: Faster generation reduces accumulated delay across long sequences of browser or application actions.
  • Live developer experiences: Voice interfaces, tab completion, streaming suggestions, and code-editing loops benefit from shorter waits between input and output.

Spend the premium where latency costs more

At current output rates, generating 100,000 tokens costs about $1 in Standard mode and $6 in Ultrafast mode, excluding input tokens. The additional $5 may be economical during a production incident or a user-blocking coding session. Background jobs, offline evaluations, and asynchronous processing can remain on Standard when completion time has little operational value.

Standard, Fast, and Ultrafast give developers three processing tiers for the same model family. Applications can now choose a tier per request based on latency targets, workload economics, and user interaction patterns.

Trending
  • No trending articles

Comments

avatar

Next Reads