Cerebras CS-4 Hits 30x Faster Inference Without Building a New Chip
Cerebras unveils CS-4 with three overclocked WSE-3 Turbo wafers, delivering 750 PFLOPS and up to 30x faster inference than GPU systems.
- Cerebras launched CS-4, a rack-scale AI system with three WSE-3 Turbo wafers delivering 750 PFLOPS
- Claims up to 30x faster inference than GPUs and 10x more throughput per watt than CS-3
- WSE-3 Turbo is not new silicon, just the WSE-3 overclocked with twice the power delivered
- New Nexus architecture uses modular wafer backpacks with 50% fewer components and 2-microsecond wafer-to-wafer latency
- Targets frontier inference: 1,000+ tokens per second on 10T-parameter models for agentic workloads
- First shipments begin this quarter; see the official blog for details
Cerebras just pulled back the curtain on CS-4, its next rack-scale AI system, and the headline numbers are aggressive: up to 30x faster inference than GPU systems and up to 10x more throughput per watt than the CS-3 it replaces. What is interesting is how they got there. Instead of shipping a brand new silicon generation, Cerebras built a new rack architecture around an overclocked version of its existing wafer, and squeezed twice the work out of the same chip.
What is actually inside the box
The Cerebras CS-4 is a rack-scale AI accelerator built from three WSE-3 Turbo wafers, delivering 750 PFLOPS of compute. It is the first system on the Nexus platform and targets ultrafast inference for very large AI models. The full spec sheet reads like a bandwidth flex:
- 750 PFLOPS of AI compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of system I/O bandwidth
- Wafer-to-wafer latency as low as two microseconds to support very large AI models
- Up to 10x more throughput per watt than CS-3
- Up to 2x faster than CS-3 and up to 30x more tokens-per-second-per-user than leading GPU solutions
Each of the three wafers is a WSE-3 Turbo. Like the WSE-3, the WSE-3T is the largest AI processor ever built, containing four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon, with 44GB of SRAM integrated directly on the wafer. The WSE-3T doubles AI compute to 250 PFLOPS per wafer and doubles memory bandwidth to 43.2 petabytes per second.
The trick: same silicon, twice the juice
Here is the part worth understanding. The WSE-3T is not a new chip. The CS-4 machines are getting an overclocked version of the current WSE-3 waferscale compute engine, with the exact same 900,000 cores and the exact same 44 GB of on-wafer SRAM, and made using the same 5 nanometer processes. If you look at the chart, you'll notice it accomplishes this using the same process tech, wafer area size, transistor count, core count, and SRAM capacity. That's because the WSE-3T isn't new silicon. Instead, Cerebras tells us it's just pushing its existing wafer scale engine harder.
The doubling in throughput comes almost entirely from power delivery and cooling. The Nexus Platform Architecture drives significant power delivery improvements. By moving power conversion 100 times closer to the processors compared to conventional GPU boards, CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3 Turbo, enabling higher operating frequencies and faster token generation. Independent estimates suggest the clock has doubled: the power and cooling technology that drives this clock speed, which is believed to have doubled to 2.8 GHz from the 1.4 GHz used in the plain vanilla WSE-3 engine, was not yet there before now.
That extra power is not free. The WSE-3 was already a hot chip at 15 kW at the wafer level and around 23 kW at the system level. This means we're probably looking at around 46 kW for each CS-4 backpack and a total system power of between 120 kW and 140 kW. This is the kind of density that only makes sense with direct liquid cooling and hyperscale data-center facilities.
Nexus: rebuilding the server around the wafer
Alongside the chip refresh, Cerebras is introducing the Nexus Platform Architecture, which is a rack-level redesign built on three modular blocks: compute, power, and I/O. The compute unit is the star of the show.
Each Wafer-Scale Backpack is a self-contained assembly that folds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours. Power delivery sits just 0.5 millimeters from the processor, roughly 100x closer than the ~50mm of conventional GPU boards.
The I/O side is equally opinionated. Based on the specs, which say "new higher speed wafer links," the wafer I/O module has six Ethernet ports running at 200 Gb/sec, which is twice the bandwidth per port as the 100 Gb/sec I/O modules used with the prior CS-1 through CS-3 systems. Crucially, wafers can be linked within and across racks without a switch, which is how Cerebras gets that two-microsecond wafer-to-wafer latency.
Why this matters for inference workloads
Cerebras has been positioning itself as the fastest inference option for frontier-scale models, and CS-4 is aimed squarely at agentic systems where latency compounds across tool calls and reasoning steps. According to the company, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters while preserving interactive decode performance.
Sean Lie, Cerebras CTO, framed the value proposition around headroom rather than perceived speed. "Being 30 times faster doesn't just make a response feel fast. It gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time," said Sean Lie, CTO and co-founder of Cerebras. "That's the difference CS-4 makes for real production workloads."
The practical workloads this targets:
- Real-time agentic pipelines that need to chain many LLM calls without users noticing the latency
- Serving frontier models (10T+ parameters) with interactive-feeling decode speeds
- Reasoning-heavy inference where token generation time dominates cost per query
- Hyperscale deployments where throughput-per-watt determines whether a workload is economical
What CS-4 is not aimed at is small-model training or general-purpose accelerated compute where a fleet of GPUs is already cost-effective. This is a purpose-built inference machine for models that are too large or too latency-sensitive for conventional clusters.
The industry context
This launch has one notable absence. While Cerebras is launching a new system and a new rack design that wraps around it with the CS-4 systems, which are codenamed "Nexus" and which lays the foundation for several more generations of machines from Cerebras, what it is not launching is a new WSE-4 compute engine as many expected. The read here is that Cerebras is decoupling silicon cadence from system cadence, betting that better power delivery and rack design can deliver a full generational jump without a new tapeout.
For anyone tracking the AI accelerator space, that is the assumption worth updating. The narrative around wafer-scale computing has always been about the chip itself. CS-4 argues that the next round of gains is coming from packaging, power, and interconnect, and that a well-engineered rack can extract another 2x from silicon that has been shipping for two years.
Availability
Pricing has not been disclosed publicly, which is standard for hardware at this scale where deals are negotiated per deployment. First CS-4 shipments are expected to begin this quarter, according to Cerebras. Access for developers will continue to flow through Cerebras Cloud, where existing WSE-3-based inference APIs are already available. If you are running inference workloads on frontier-scale models and hitting latency or throughput ceilings on GPU clusters, this is the system to benchmark against.