Cerebras Runs Google's Gemma 4 31B at 1,800 Tokens per Second

Gemma 4 31B hits Cerebras public preview at 1,800+ tokens/sec, marking the platform's first multimodal model and first Google DeepMind collaboration

·
·
Read6 min
TypeNews
TopicLlms · Gpus
  • Public preview live: Gemma 4 31B is now available on Cerebras Inference Cloud, the platform's first multimodal model.
  • 1,800+ tokens/sec: Cerebras claims it's the world's fastest multimodal model, running 15-18x faster than Claude Haiku at comparable quality.
  • First Google DeepMind model on Cerebras: Marks a new partnership and the platform's first vision-capable model, with more multimodal models planned.
  • Full vision capabilities: Accepts screenshots, documents, charts, and UI states; supports native function calling, 256K context, and reasoning mode.
  • Priced competitively: ~$0.12/M input and $0.35/M output tokens via OpenRouter; Cerebras-specific pricing coming soon.
  • Key limitation: Images must be base64-encoded (no external URLs), max 5 images/request, and reasoning mode is off by default.

Cerebras just crossed two milestones at once. Gemma 4 31B is the first Google DeepMind model they have brought to the platform, and the first to let developers feed images into a model running at wafer-scale speed. The result is something the inference market hasn't seen before: a capable multimodal model that responds fast enough to feel interactive.

What just shipped

Gemma 4 31B is available today on the Cerebras Inference Cloud in public preview. It runs at over 1,800 tokens per second on Cerebras Inference, making it the world's fastest multimodal model. For context, Claude Haiku runs at roughly 100 tokens per second, making this a 15x speedup against one of the most popular production-grade models on the market, at roughly equivalent quality.

Gemma 4 31B is comparable to Claude Haiku 4.5 in intelligence, scoring 29 and 30 respectively on the Artificial Analysis Intelligence Index. The key difference is that Gemma 4 is open-weight under Apache 2.0, and on Cerebras it runs 18x faster than Haiku.

The model itself

Gemma 4 is Google DeepMind's most intelligent open model family, built from Gemini 3 research. The 31B is the flagship and most capable model in the family, a dense multimodal model built for quality and efficiency rather than raw parameter count. Dense models achieve high model intelligence without the large memory footprint of MoE (Mixture-of-Experts) models. MoE is an architecture where only a subset of the model's parameters activates per request, trading some quality for speed and lower memory use.

The 31B model currently ranks as the #3 open model in the world on the industry-standard Arena AI text leaderboard, and outcompetes models 20x its size. On reasoning benchmarks, it delivers 89.2% on AIME 2026, 85.2% on MMLU Pro, 80.0% on LiveCodeBench v6, and 84.3% on GPQA Diamond.

The model's capabilities go well beyond text:

  • Image understanding: object detection, document/PDF parsing, screen and UI understanding, chart comprehension, OCR including multilingual, handwriting recognition, and pointing.
  • Interleaved multimodal input: freely mix text and images in any order within a single prompt.
  • Function calling: native support for structured tool use, enabling agentic workflows.
  • A context window of 256K tokens, determining how much text it can process in a single interaction.
  • Multilingual: out-of-the-box support for 35+ languages, pre-trained on 140+ languages.

Google builds the Gemma models using knowledge distillation from Gemini, essentially training the smaller model to mimic the reasoning patterns of a much larger one. This is a key reason Gemma models punch above their parameter count.

Why speed changes everything here

Raw throughput numbers are easy to dismiss until you think about what multimodal agents actually do. Multimodal and agentic loops rarely call a model once: they inspect a visual input, reason over it, produce structured output, call tools, check the result, and try again. At conventional speeds those loops are too slow to provide real-time input. At over 1,800 TPS, the application and user work in lockstep.

This is the real unlock. A pipeline that takes a screenshot, extracts structured data, calls an API, and verifies the result might need 5-10 model calls. At 100 tokens/sec, that's a multi-second wait per step. At 1,800 tokens/sec, the whole loop feels instant. Gemma 4 is the first model on Cerebras to support image understanding, enabling workflows combining text with images: screenshots, charts, UI states, scanned pages, forms, diagrams. It also unlocks computer use and robotics applications.

Practical use cases

The Cerebras team highlights three concrete workflows where this combination of speed and vision pays off:

  • Screenshot insight: Feed a dense dashboard screenshot or document page, get structured output identifying what matters, in real time rather than after a wait.
  • Long-context summarization: Hand it a research report or technical brief and get a decision-ready summary fast enough to read, react, and re-query in a single sitting.
  • Screenshot to patch: Take a broken UI screenshot, the source code, and the console error, and get back a minimal patch with verification checks.

The model excels at multimodal reasoning across screenshots, documents, diagrams, and design assets, making it ideal for visual agentic workflows, image-aware copilots, and teams migrating from closed multimodal APIs to an open model.

Limitations to know

A few constraints are worth knowing before you build around this:

  • Image inputs are only supported in the Chat Completions endpoint, not the Completions endpoint.
  • The free tier caps context at 65k tokens (131k on paid), with a maximum of 5 images per request and 10 MB total image payload.
  • External image URLs are not supported; you must pass images as base64-encoded PNG or JPEG data URIs.
  • Reasoning mode is disabled by default; enable it via the reasoning_effort parameter. The raw and hidden reasoning formats are not supported.
  • Gemma 4 is not at the absolute frontier. GPT-4o, Claude 3.7, and Gemini 2.0 Ultra still outperform it on the most complex reasoning tasks. If you need state-of-the-art performance on hard agentic benchmarks, a closed model still has the edge.

How to access it and what it costs

The model is live now on the Cerebras Inference Cloud in public preview. The free tier allows 5 requests/min, 30k input tokens/min, and 1M tokens/day. Pay-as-you-go bumps that to 300 requests/min and 500k input tokens/min. Cerebras has not yet published specific per-token pricing for Gemma 4 31B on their platform, though on OpenRouter the model runs at $0.12 per million input tokens and $0.35 per million output tokens, which is a fraction of what comparable closed models cost.

To call the model, it's a standard OpenAI-compatible API call with the model ID gemma-4-31b:

python
from cerebras.cloud.sdk import Cerebras
import base64
client = Cerebras()
# Encode your image as base64
with open("screenshot.png", "rb") as f:
    img_b64 = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
    model="gemma-4-31b",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url",
             "image_url": {"url": f"data:image/png;base64,{img_b64}"}},
            {"type": "text",
             "text": "Extract all KPIs from this dashboard as JSON"}
        ]
    }]
)
print(response.choices[0].message.content)

The bigger picture

Bringing vision to wafer-scale hardware is a milestone for the platform. Multimodal support starts with Gemma 4, and Cerebras will extend it to additional models going forward. This is the first signal that Cerebras is moving beyond text-only inference, and the Google DeepMind partnership opens the door to more joint releases. For teams currently paying for Claude Haiku or GPT-4o-mini to power visual pipelines, an open-weight alternative running 15-18x faster at lower cost is a meaningful shift in what's economically viable to build.

Trending
  • No trending articles

Comments

avatar

Next Reads