OpenAI's GPT-5.6 Sol Hits 750 Tokens per Second on Cerebras Hardware

OpenAI's Ultrafast tier runs GPT-5.6 Sol at 750 tokens per second on Cerebras hardware, up to 14x faster than standard processing, unlocking real-time frontier AI for enterprise workflows.

·
·
AuthorOpenAI
Read2 min
TopicLlms · Api
SubtopicLong Context
  • OpenAI previews Ultrafast mode: GPT-5.6 Sol runs at up to 750 tokens/sec, 14x faster than standard processing, via the OpenAI API.
  • Powered by Cerebras WSE-3: The wafer-scale chip eliminates GPU memory bottlenecks; OpenAI has a $20B+ multi-year compute deal with Cerebras.
  • Speed vs. predecessor: 7-10x faster than GPT-5.5 XHigh (70-100 TPS) and ~5x faster than typical production GPU deployments (~150 TPS).
  • Key use cases: Real-time voice, incident response, financial research, commerce, and live experimentation where latency has direct business value.
  • Limited access for now: Available to a select group of API customers; businesses can sign up for access updates as capacity expands.
  • Competitive gap closed: Groq offers comparable speeds but only for open-source models; Ultrafast is the first frontier proprietary model at this speed tier.

OpenAI just drew a new line in the AI inference race. The company is previewing Ultrafast mode, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, 14 times faster than standard processing. The catch: it's launching in limited preview to a select group of API customers, with broader access gated by capacity. But the implications are significant enough to pay attention to now.

Speed as a product feature

For most of AI's recent history, getting faster inference meant accepting a weaker model. You could have speed, or you could have intelligence. Ultrafast is OpenAI's argument that you no longer have to choose. Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.

To understand what 750 tokens per second actually means in practice, consider the baseline. GPT-5.6 Sol runs at up to 750 tokens a second on Cerebras hardware, roughly 5x the approximately 150 tokens a second most production models deliver today. For a single chatbot exchange, that difference might feel cosmetic. But for agentic workflows, it compounds fast. For an enterprise agent that chains 30 or 40 model calls to finish one back-office task, that 5x compounds into the difference between a workflow that finishes in seconds and one that finishes in minutes.

The predecessor model, GPT-5.5 XHigh, ran at roughly 70 to 100 tokens per second. GPT-5.6 Sol is 7 to 10 times faster than its predecessor at the high end. That is not an incremental improvement. That is a different product category.

The hardware story: why Cerebras

The speed comes from an unusual piece of silicon. While traditional GPUs stitch together dozens of discrete chips, Cerebras builds one giant processor, the Wafer-Scale Engine, that keeps all the model weights local. No memory bottlenecks, no interconnect latency. Just raw, uninterrupted compute flow.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves