Gimlet Cloud Adds Cerebras to Hit 3,000 Tokens per Second for AI Agents
Gimlet Labs is deploying 100 megawatts of Cerebras wafer-scale systems inside its multisilicon inference cloud, promising up to 3,000 tokens per second.
- Gimlet Labs and Cerebras announced a partnership deploying 100 MW of wafer-scale inference capacity.
- First datacenter comes online later this year, targeting up to 3,000 tokens per second.
- Gimlet Cloud disaggregates inference phases across GPUs and Cerebras systems via one API.
- Company claims 3-10x gains in interactivity or throughput-per-kilowatt versus homogeneous GPU clusters.
- Follows $300M Series B led by a16z at a $3B valuation weeks earlier.
- Gimlet is also a launch partner for the next-generation Cerebras CS-4 arriving in 2027.
Gimlet brings Cerebras into its 100-megawatt inference plan
San Francisco startup Gimlet Labs, which builds software to route AI inference across several processor types, will add Cerebras wafer-scale systems to Gimlet Cloud under a strategic partnership. The companies plan 100 megawatts of capacity, with the first data center in the rollout expected online later in 2026. Gimlet is targeting peak output of up to 3,000 tokens per second for AI agents and other latency-sensitive applications.
Private deployments are already serving traffic through the integration, according to the joint announcement. Pricing, regions, supported models, live capacity and a public launch date remain unpublished. The 100-megawatt figure describes planned capacity. The 3,000-token figure is a peak target whose test conditions have yet to be released.
Where agent latency multiplies
AI agents often chain dozens of model calls, with each call depending on the previous result. Slow token generation therefore limits how many sources an agent can inspect, how many approaches it can test and how quickly it can finish a task. With a fixed token count, raising generation speed from 100 to 3,000 tokens per second cuts the generation portion of a 10-minute workload to about 20 seconds. That estimate excludes prompt processing, queueing, network delay and tool execution.
One request, several chips
Large language model inference has two main phases. Prefill processes the prompt in parallel and creates the attention state needed for generation. Decode produces tokens sequentially and repeatedly reads model weights and that stored state. Prefill tends to reward raw parallel compute, while decode at latency-sensitive batch sizes often depends on memory bandwidth.
| Workload phase | Primary work | Likely hardware role |
|---|---|---|
| Prefill | Process prompt tokens and build the key-value cache | GPUs optimized for parallel computation |
| Decode | Generate tokens sequentially while reading weights and cached state | Cerebras systems and other accelerators with high memory bandwidth |
| Specialized kernels | Run attention, feed-forward and speculative-decoding operations | GPUs, CPUs or accelerators selected by Gimlet’s scheduler |
Moving a request between processor types creates its own costs. The scheduler may need to transfer the key-value cache, which stores the attention state created during prefill, across systems with different memory layouts. Serialization, interconnect bandwidth, queueing and compiler support can erase hardware gains if orchestration is slow. Gimlet’s performance claims therefore depend on the entire serving stack, including how efficiently it moves state between chips.
SRAM changes the decode equation
Cerebras’ WSE-3 places compute cores and SRAM across a single wafer. The company lists 4 trillion transistors, 900,000 AI cores, 44 GB of on-chip SRAM and 21 petabytes per second of memory bandwidth. Larger models can be partitioned across multiple systems. Keeping memory close to compute reduces data movement during autoregressive generation, where the system repeatedly reads weights to produce one token at a time.
Cerebras compares the WSE-3’s on-chip memory bandwidth with the Nvidia H100 at a ratio of roughly 7,000 to one. End-to-end inference includes additional constraints such as inter-system communication, batching, prompt processing and software overhead. Artificial Analysis reported output above 2,700 tokens per second for gpt-oss-120B on a Cerebras CS-3 endpoint, compared with about 900 tokens per second on a Blackwell B200 endpoint. Endpoint results reflect the provider’s full stack, and comparisons can vary with precision, batching, prompt length and output length.
Peak speed meets production load
GPU serving systems commonly raise aggregate throughput by batching more requests, which can increase latency for each user. A Pareto frontier describes the best combinations of those two measures: per-user speed and total system throughput. Gimlet says its multi-silicon design shifts that frontier in two ways:
- Interactivity: Three to ten times higher per-user generation speed at a fixed throughput-efficiency target.
- Power efficiency: Three to ten times more throughput per kilowatt at a fixed per-user speed.
Published materials omit the methodology behind those ranges, including batch sizes, concurrency, model precision, prompt lengths, time to first token and tail latency. They also leave unclear whether the 3,000-token target represents a single stream or aggregate output. Higher throughput per kilowatt could reduce serving costs, but developers still need pricing and sustained-load measurements to calculate cost per completed task.
The abstraction developers touch
Applications send requests through one Gimlet inference API, and the orchestrator assigns each phase to available hardware. Cerebras will become a native accelerator within the platform. Gimlet has also signed on as a launch partner for the next-generation CS-4, which is expected to reach customers in 2027.
The published announcement omits code examples, an API compatibility specification, service-level commitments, quotas, regional availability and a model coverage matrix. Those details will determine whether applications can move between hardware backends without changes to request schemas, streaming behavior, tool calls or token accounting.
Capital behind the buildout
The partnership followed Gimlet’s Series B by about three weeks. The company raised $300 million at a $3 billion valuation, six months after an $80 million Series A, bringing reported funding to $392 million. Andreessen Horowitz led the round, with participation from Sapphire Ventures, M12 and Arm. Sapphire says Gimlet is scaling toward hundreds of megawatts of managed heterogeneous infrastructure and has billions of dollars in contracted revenue, though contract details remain private.
Deploying 100 megawatts requires hardware reservations, power agreements, networking and data-center construction. The recent financing provides context for that capital-intensive plan, while the rollout schedule offers only one public milestone: the first data center expected later in 2026.
Gimlet’s founders previously built Pixie Labs, a Kubernetes observability company based on eBPF, Linux’s in-kernel programmability layer. New Relic acquired Pixie in 2020. CEO Zain Asgar is an adjunct professor at Stanford and previously worked as a GPU architect at Nvidia and an engineering lead at Google AI. Gimlet now supports hardware from Nvidia, AMD, Intel, Arm, Cerebras and d-Matrix, reflecting its goal of operating as a vendor-neutral inference layer.
What production teams should test
- Latency under load: Measure time to first token, output speed, complete-task latency and p50 and p95 results at expected concurrency.
- Model fidelity: Confirm checkpoint versions, quantization, context limits, sampling controls and output quality across backends.
- API behavior: Exercise streaming, structured output, tool calling, cancellation, retries and token accounting.
- Routing and recovery: Trigger accelerator saturation and failures to observe fallback behavior, latency changes and error handling.
- Observability: Verify that logs and metrics identify the hardware path, queue time, prefill time and decode time for each request.
- Data controls: Confirm processing regions, retention policies, encryption, isolation and compliance commitments.
- Economics: Compare price per completed task at realistic prompt lengths, output lengths and concurrency levels.
The Cerebras integration gives Gimlet a concrete deployment path for routing inference phases across unlike processors. The companies report live private traffic. Public evaluation still requires sustained concurrency data, tail latency, reliability, model coverage and pricing. The announced 100 megawatts remain planned capacity, and 3,000 tokens per second remains a peak performance target.