Cerebras Runs Google DeepMind's Gemma 4 at 1,500 Tokens per Second

Gemma 4 31B lands on Cerebras Inference at 1,500 tokens/sec — 15x faster than Claude Haiku — with a 24-hour hackathon kicking off this Sunday

·
·
Cerebras Runs Google DeepMind's Gemma 4 at 1,500 Tokens per Second
Read5 min
TypeNews
TopicLlms · Gpus
  • Speed milestone: Gemma 4 31B runs at 1,500+ tokens/sec on Cerebras — 15x faster than Claude Haiku at comparable quality.
  • First multimodal on Cerebras: Gemma 4 is the first image-capable model on the platform, enabling real-time vision + agentic loops.
  • 24-hour hackathon this Sunday: Register on Luma for early access, $5K in prizes, and a Q&A with the Gemma 4 research team.
  • Model specs: 30.7B dense params, 256K context window, Apache 2.0, supports OCR, PDF parsing, UI understanding, function calling, and 140+ languages.
  • Private preview now, GA end of June: Available via Cerebras Inference; weights free on Hugging Face.
  • Best for: Agentic loops, document/screenshot workflows, long-context summarization, and coding tasks where inference speed is the bottleneck.

Cerebras and Google DeepMind just crossed a milestone together: Gemma 4 31B is now running on Cerebras Inference, making it the first multimodal model on the platform and the first Google DeepMind model Cerebras has ever hosted. To celebrate, they're running a 24-hour virtual hackathon this Sunday with $5,000 in prizes and early API access for participants.

Speed that changes what's possible

The headline number is hard to ignore. Cerebras runs Gemma 4 at over 1,500 output tokens per second. By comparison, Claude Haiku runs at roughly 100 tokens per second. That's a 15x speedup against one of the most popular production-grade models on the market, at roughly equivalent quality.

But raw speed only matters if it changes what you can build. Here it does. Multimodal and agentic loops rarely call a model once: they inspect a visual input, reason over it, produce structured output, call tools, check the result, and try again. At 100 tokens per second those loops are too slow to provide real-time input. At 1,500 TPS, the application and user can work together at the same time. Front-end iteration feels near-instant, document and screenshot workflows return in a fraction of the time, and developers can fit more verification and more retries into the same product.

What Gemma 4 31B actually is

Gemma 4 31B is the flagship of Google DeepMind's open-weight Gemma family , a dense model (meaning all parameters are active on every forward pass, unlike sparse Mixture-of-Experts models) built for quality and efficiency rather than raw parameter count. Dense models achieve high model intelligence without the large memory footprint of MoE models. Gemma 4 hits a sweet spot: strong enough for serious work, efficient to serve, and open enough to build around without vendor lock-in.

It's a 30.7B parameter multimodal model supporting text, images, and video with a 256K context window, native thinking mode, function calling, and 140+ languages , released under Apache 2.0. The 256K context window is double what Gemma 3 offered, and the Apache 2.0 license means you can use it commercially without restrictions.

On benchmarks, on the Artificial Analysis Intelligence Index, Gemma 4 31B scores 29 , effectively matching Claude Haiku at 30. The difference is that Gemma 4 is open-weight under Apache 2.0, and on Cerebras it runs an order of magnitude faster.

The multimodal capabilities worth knowing

Gemma 4 is the first model on Cerebras to support image understanding. It enables workflows combining text with images , screenshots, charts, UI states, scanned pages, forms, diagrams. It also unlocks computer use and robotics applications.

The image understanding feature set is broad:

  • Object detection, document/PDF parsing, screen and UI understanding, chart comprehension, OCR (including multilingual), handwriting recognition, and pointing.
  • Video understanding by analyzing sequences of frames, with the ability to freely mix text and images in any order within a single prompt.
  • Native function calling for structured tool use, enabling agentic workflows.
  • Configurable thinking modes , all models in the family are designed as highly capable reasoners.

One architectural detail worth understanding: Gemma 4 uses hybrid attention , alternating between local sliding window attention (1024 tokens) and global full attention, balancing efficiency with long-range coherence. This is what lets it handle 256K context without blowing up memory. The vision encoder is a ~550M parameter component supporting variable aspect ratios and configurable token budgets (70 to 1120 tokens per image), making it useful for both quick captioning and dense OCR.

What to build with it

The Cerebras team is explicit about where this combination shines. The practical use cases they highlight:

  • Screenshot insight: Feed a dense dashboard screenshot or document page, get structured output identifying what matters , in real time.
  • Long-context summarization: Hand it a research report or technical brief and re-query in a single sitting.
  • Screenshot to patch: Give it a broken UI screenshot, the source code, and the console error , it returns a minimal patch and the checks to verify it.

If you are building multimodal reasoning, document understanding, fast summarization, or targeted coding workflows and inference speed is the bottleneck, this is the platform to watch.

The hackathon: early access and $5K in prizes

Cerebras and Google DeepMind are running a Gemma 4 24-Hour Hackathon this Sunday, June 28 at 10:00 AM PT through June 29 at 10:00 AM PT. It's fully virtual, hosted on the Cerebras Discord. The event kicks off with a 30-minute intro and Q&A with the Gemma 4 research team.

Prize structure:

  • Best Enterprise Use Case , $2,500
  • Best Inference Speed Demo , $2,500

Participants get early access to Gemma 4 on Cerebras before general availability. Winners are also invited to showcase their demos at Cafe Compute SF the next evening. You can register on Luma and submit as many projects as you want via the Discord #showcase channel.

Availability and pricing

Gemma 4 is now in private preview on Cerebras Inference, with general availability later this month. For reference, the model is also available on other inference providers: on Together AI, it runs at $0.20 per million input tokens and $0.50 per million output tokens. Cerebras pricing is available on their pricing page. The model weights themselves are free to download from Hugging Face under Apache 2.0.

Bringing vision to wafer-scale hardware is a milestone for the platform. Multimodal support starts with Gemma 4, and Cerebras will extend it to additional models going forward. The combination of image understanding and wafer-scale speed is what unlocks new product experiences: a model that can see a dashboard, reason over it, return structured output, and act on it fast enough to keep a human or an agent in the loop.

Trending
  • No trending articles

Comments

avatar

Next Reads