Cerebras Nearly Quadruples Cloud Revenue but Drops 15% After Hours
Cerebras Q2 cloud revenue surges 281%, beats estimates, lands OpenAI and AMD deals — but stock drops 15% after hours

- Cloud revenue nearly quadrupled: GAAP cloud revenue hit $126M (+281% YoY); core total revenue reached $210M (+103% YoY).
- Stock dropped ~15% after hours despite beating EPS estimates, as GAAP revenue of $180M missed the $194M analyst consensus.
- OpenAI partnership deepened: Cerebras is now the launch partner for GPT-5.6 Sol, serving it at 750 tokens per second.
- Disaggregated inference with AMD and AWS: AMD Helios handles prefill, Cerebras WSE handles decode — claiming 5x throughput gains; production in Q4 2026, AWS Bedrock in Q1 2027.
- Massive backlog and raised guidance: $25.4B in remaining performance obligations; full-year 2026 core revenue guidance raised to $880–$890M; plans to triple revenue in 2027.
- Supply chain moat: Cerebras avoids HBM, CoWoS, and 3nm — the three most constrained components in AI chip manufacturing — giving it a structural scaling advantage.
Cerebras Systems just posted its second quarterly earnings report as a public company, and the numbers are hard to ignore. Cloud and other services revenue hit a record $126 million on a GAAP basis, up 281% year-over-year, while core cloud revenue reached $127.7 million, up 287%. The cloud business has nearly quadrupled in a single year. The market shrugged, at least initially.
Following the earnings release, CBRS declined roughly 15% in after-hours trading, even as the company beat on EPS, reporting an adjusted loss of 5 cents versus 17 cents expected. The culprit was GAAP revenue of $180 million, below the $194 million analysts had penciled in, with the gap driven largely by how Cerebras accounts for its non-standard revenue structure.
The numbers behind the headline
To parse Cerebras's financials, you need to know the difference between its GAAP and "core" (non-GAAP) numbers. GAAP total revenue came in at $180.1 million, up 74% year-over-year, while core total revenue, which strips out pass-through data center costs and adds back non-cash customer warrant amortization, reached $209.9 million, up 103%. The company's guidance and analyst expectations were set against the core number, which is why the beat/miss picture confuses at first glance.
- Record GAAP cloud revenue: $126.0M (+281% YoY); core cloud revenue: $127.7M (+287% YoY)
- Core gross margin: 41%, an improvement of roughly 940 basis points from Q2 2025
- Core EPS beat: reported loss of $0.04 per share vs. estimates of -$0.17
- Remaining performance obligations (RPO): $25.4 billion as of June 30, 2026
- Data center capacity under contract: over 600 MW; manufacturing capacity scaling more than 10x in 2026
- Full-year 2026 core revenue guidance raised to $880–$890 million, from a prior range of $855–$865 million
CEO Andrew Feldman summarized it simply: "Core revenue more than doubled to $210 million, and our cloud business nearly quadrupled year-over-year." CFO Bob Komin added that the company plans to more than triple revenue in 2027. That target is aggressive, but it is backstopped by $25.4 billion in already-contracted future revenue.
Why speed is the product
Cerebras's core thesis is that inference latency doubles as a product differentiator, unlocking use cases that slower hardware cannot serve. The company's Wafer-Scale Engine (WSE-3) is the silicon behind that claim. While conventional AI accelerators like Nvidia's H100 are cut from a wafer into individual dies of around 814 mm², the WSE-3 uses an entire TSMC 5nm wafer as a single chip, measuring 46,225 mm², roughly 57x the area of an H100.
The critical advantage is memory bandwidth and latency. Because the WSE-3 holds 44 GB of on-chip SRAM directly accessible to 900,000 cores, rather than routing through HBM stacks connected via PCIe or NVLink, it eliminates the memory bottleneck that limits GPU-based inference at scale. That translates directly to tokens per second, which is what customers pay for. This quarter, Cerebras enabled support for OpenAI's GPT-5.6 Sol at 750 tokens per second.
There is a supply chain angle too. Through its wafer-scale architecture, Cerebras avoids many components currently in short supply. It does not use HBM memory, CoWoS packaging, or 3nm fabrication technology, all of which are supply-limited today. Memory is baked into the wafers, so no HBM. The wafer is never cut into pieces that need re-stitching, so no CoWoS. And the design targets the older, less-crowded 5nm node. While competitors fight over scarce technologies, Cerebras largely sidesteps the traffic jam.
The disaggregated inference bet
One of the most technically interesting announcements this quarter is Cerebras's push into what the industry calls disaggregated inference, the practice of splitting an AI inference job across two specialized chips, each handling the phase it is best at.
Running a large language model has two distinct phases. Prefill, reading the prompt and its context, is compute-heavy and rewards throughput, which is what a GPU rack like AMD Helios does well. Decode, writing each new token, is memory-bandwidth-heavy and rewards ultra-low latency, which is where a single giant Cerebras wafer shines. Forcing one chip to do both means paying for capacity the other phase wastes. Splitting the two lets each phase run on hardware built for it.
AMD and Cerebras announced a technical partnership combining AMD's Helios rackscale solution with Cerebras's Wafer-Scale Engine to deliver a disaggregated AI inference system. The companies expect the platform to deliver up to 5x higher tokens per second per watt by assigning different portions of an inference workload to architectures optimized for them. The AMD solution is expected to enter production in Q4 2026.
Cerebras is running the same playbook with AWS, using Trainium for prefill, CS-3 systems for decode, and Amazon's Elastic Fabric Adapter to move data between them. The company expects to bring the same 5x throughput benefits to Amazon Bedrock in Q1 2027.
A caveat on the numbers: the 5x tokens-per-watt claim is a model against a Cerebras-only setup, not a measured benchmark against Nvidia. Cerebras is exceptional at single-user speed, but the company does not publish the kind of high-concurrency aggregate throughput curves that would prove it wins on tokens per second per dollar at scale. The architecture is genuinely novel, and the performance claims need independent validation in production.
A customer list that signals market expansion
Beyond the OpenAI anchor deal, the customer additions this quarter show where fast inference is finding traction. Cerebras signed new cloud capacity agreements with leading AI coding companies including Cognition and Lovable. Fast inference also serves as the foundation for agentic flows in industries ranging from finance to life sciences, with customers including Block, Figma, AlphaSense, and GSK.
The most unexpected vertical is security. Cerebras pioneered a new segment of the security market with CrowdStrike, where fast inference enables inline security using LLMs for a large portion of enterprise traffic. Real-time threat detection with an LLM in the critical path requires latency that GPU clusters cannot deliver at scale, which is exactly the kind of workload Cerebras's architecture targets.
Why the stock dropped anyway
Cerebras raised its full-year guidance in its second earnings report following its May IPO, the largest semiconductor IPO on record, which raised $6.4 billion. Despite beating on EPS and raising guidance, shares declined roughly 11–15% in after-hours trading.
Cerebras's current Price-to-Sales ratio sits at an elevated 291.5x, well above its historical median of roughly 238.6x and the semiconductor industry average, signaling steep expectations for future growth. At that valuation, even strong results can disappoint if the top-line GAAP number misses consensus. The $450 million GAAP net loss also draws attention, though most of it ties to stock-compensation costs of $386.6 million, a non-cash item that inflates the GAAP figure dramatically post-IPO.
Cloud and services now represent roughly 70% of quarterly revenue, reflecting a channel shift toward hosted inference. That mix shift is what investors want to see: recurring, higher-margin cloud revenue replacing lumpy hardware sales. Margin trajectory is moving in the right direction, even if absolute numbers remain negative. To satisfy contracted demand before its own data-center infrastructure comes online, Cerebras has been temporarily renting back systems from an existing customer, a cost that reduces cloud and services gross margin by roughly 10–15 percentage points until rented capacity is replaced with Cerebras-controlled deployments.
What comes next
The near-term roadmap is dense. The AMD disaggregated inference system hits production in Q4 2026. The AWS Bedrock integration follows in Q1 2027. Manufacturing capacity is scaling more than 10x this year across contract manufacturers Flex, Sanmina, and Rocket EMS. The company has also secured TSMC wafer supply for continued growth.
The bigger question is whether Cerebras can convert its $25.4 billion RPO backlog into recognized revenue fast enough to justify its valuation. The OpenAI commitment and AWS partnership validate the technology and are expected to support substantial capacity expansion over the next several years. With cloud and services at roughly 70% of quarterly revenue and a clear path to disaggregated inference at hyperscale, the architecture thesis is proving out. Execution risk now sits in the buildout: data centers, manufacturing lines, and a supply chain that needs to scale 10x in months, not years.