OpenAI's GPT-5.6 Sol Hits 750 Tokens per Second on Cerebras Hardware
OpenAI's Ultrafast tier runs GPT-5.6 Sol at 750 tokens per second on Cerebras hardware, up to 14x faster than standard processing, unlocking real-time frontier AI for enterprise workflows.
- OpenAI previews Ultrafast mode: GPT-5.6 Sol runs at up to 750 tokens/sec, 14x faster than standard processing, via the OpenAI API.
- Powered by Cerebras WSE-3: The wafer-scale chip eliminates GPU memory bottlenecks; OpenAI has a $20B+ multi-year compute deal with Cerebras.
- Speed vs. predecessor: 7-10x faster than GPT-5.5 XHigh (70-100 TPS) and ~5x faster than typical production GPU deployments (~150 TPS).
- Key use cases: Real-time voice, incident response, financial research, commerce, and live experimentation where latency has direct business value.
- Limited access for now: Available to a select group of API customers; businesses can sign up for access updates as capacity expands.
- Competitive gap closed: Groq offers comparable speeds but only for open-source models; Ultrafast is the first frontier proprietary model at this speed tier.
OpenAI just drew a new line in the AI inference race. The company is previewing Ultrafast mode, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, 14 times faster than standard processing. The catch: it's launching in limited preview to a select group of API customers, with broader access gated by capacity. But the implications are significant enough to pay attention to now.
Speed as a product feature
For most of AI's recent history, getting faster inference meant accepting a weaker model. You could have speed, or you could have intelligence. Ultrafast is OpenAI's argument that you no longer have to choose. Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.
To understand what 750 tokens per second actually means in practice, consider the baseline. GPT-5.6 Sol runs at up to 750 tokens a second on Cerebras hardware, roughly 5x the approximately 150 tokens a second most production models deliver today. For a single chatbot exchange, that difference might feel cosmetic. But for agentic workflows, it compounds fast. For an enterprise agent that chains 30 or 40 model calls to finish one back-office task, that 5x compounds into the difference between a workflow that finishes in seconds and one that finishes in minutes.
The predecessor model, GPT-5.5 XHigh, ran at roughly 70 to 100 tokens per second. GPT-5.6 Sol is 7 to 10 times faster than its predecessor at the high end. That is not an incremental improvement. That is a different product category.
The hardware story: why Cerebras
The speed comes from an unusual piece of silicon. While traditional GPUs stitch together dozens of discrete chips, Cerebras builds one giant processor, the Wafer-Scale Engine, that keeps all the model weights local. No memory bottlenecks, no interconnect latency. Just raw, uninterrupted compute flow.
The Cerebras WSE-3 is not a conventional chip. Traditional GPUs are constrained by the reticle limit of semiconductor lithography, which restricts chip size to roughly 800 square millimeters. Cerebras asked what would happen if you used the entire silicon wafer, all 46,225 square millimeters, as a single chip. The result is a processor that is 57 times larger than Nvidia's H100 GPU, packing 4 trillion transistors, 900,000 AI-optimized cores, and 44 GB of on-chip SRAM memory with a bandwidth of 21 petabytes per second.
That architecture is what makes 750 tokens per second possible. The architecture eliminates the off-chip memory bottleneck that plagues GPU-based systems. In a traditional GPU cluster, data must travel between the GPU's on-chip cache, external HBM memory, and across network interconnects between servers. Each hop adds latency and consumes power.
A partnership with serious money behind it
This is not a casual vendor relationship. On January 14, 2026, Cerebras and OpenAI disclosed an initial inference deal worth more than $10 billion, under which OpenAI would deploy 750 megawatts of Cerebras CS-3 capacity to power inference workloads behind ChatGPT and the OpenAI API. That deal subsequently expanded. The relationship grew into a multi-year master agreement valued at more than $20 billion, with the same 750 MW deployed through 2028 and options for OpenAI to acquire up to an additional 1.25 gigawatts of capacity between 2029 and 2030.
The financial structure goes deeper still. As part of the arrangement, OpenAI advanced Cerebras a $1 billion working capital facility at 6% interest secured by warrants exercisable for up to 33.4 million Cerebras shares, a structure that could give OpenAI roughly an 11% equity stake in Cerebras if all warrants vested. OpenAI is not just a customer here. It is a strategic investor with skin in the game.
For Cerebras, the timing of this announcement is pointed. The startup officially joined the NASDAQ under the ticker CBRS, having raised $5.5 billion in the process. Shares skyrocketed nearly 70 percent on the first day of trading, as investors poured their money into a new way to play the AI boom. Having OpenAI's flagship model running on your hardware is the best possible reference customer for a newly public chip company.
What it unlocks
OpenAI has identified five categories where Ultrafast creates the most leverage:
- Incident response: Analyze logs, traces, and code changes to identify the cause of an outage while the system is still failing, not after.
- Financial research and security: Assess transactions and identify suspicious activity while market conditions are still in motion.
- Real-time voice and customer support: Resolve multi-step issues without interrupting the conversation flow.
- Commerce: Answer product questions, check inventory, and resolve checkout issues before a shopper abandons their cart.
- Live research and experimentation: Turn research that previously took an overnight run into an interactive working session, letting teams test an idea, examine the results, adjust their approach, and run another experiment without breaking their flow.
Early customers are already seeing the difference. Podium's product lead for Voice AI noted that "the speed completely changes the call experience for the more complex work." Rogo's Applied AI team observed that "speed doesn't just make the product feel better. It changes what people can realistically use it for."
Where it fits in the competitive landscape
Ultrafast enters a market where speed-focused inference already exists, but with a key constraint. Groq, the other major specialized inference provider, runs 4 to 12 times faster than OpenAI but only serves open-source models. Groq's LPU chips deliver token generation speeds of 500 to 1,000+ tokens per second, 5 to 14 times faster than GPU-based providers. The problem is that Groq cannot run GPT-5.6 Sol. Ultrafast is the first time frontier-class proprietary intelligence has been available at this speed tier.
That distinction matters enormously for agentic use cases. GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series, suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving. Running that model at 750 tokens per second is a different proposition than running a capable-but-smaller open-source model at the same speed.
The standard GPT-5.6 Sol API is priced at $5.00 input and $30.00 output per million tokens , with no pricing yet announced for the Ultrafast tier specifically. Given the hardware costs involved, expect a meaningful premium over standard rates when it becomes broadly available.
The capacity constraint is real
The biggest caveat is availability. OpenAI initially described the Cerebras-backed version of GPT-5.6 Sol as a limited rollout for selected customers while capacity expanded. A deployment that dedicates dozens of wafer-scale systems to each model replica would be expensive, capacity constrained, and difficult to scale instantly.
This is not a soft launch. It is a genuine infrastructure constraint. The WSE-3 is physically enormous, and deploying enough of them to serve broad API traffic takes time and capital. OpenAI is using the preview period to study which workloads benefit most, which will inform how they prioritize capacity as it grows.
For teams that need it now, the path is to sign up for access updates on OpenAI's site. For everyone else, the more important signal is what this preview reveals about the direction of frontier AI: intelligence and speed are converging, and the workflows that benefit most are the ones that chain many model calls together. If your product does that, Ultrafast is worth watching closely.