Google Boosts Gemini Nano on Pixel 50% Faster Without Retraining

Google retrofits Multi-Token Prediction onto frozen Gemini Nano v3, delivering 50%+ faster on-device inference on Pixel 9 and 10 without touching the base model

·
·
Google Boosts Gemini Nano on Pixel 50% Faster Without Retraining
  • Google retrofits Multi-Token Prediction (MTP) onto frozen Gemini Nano v3 without retraining the base model, now live on Pixel 9 and 10.
  • The MTP head cross-attends to the main model's KV cache, saving 130MB RAM per instance vs. a standalone drafter model.
  • Production speedups of 50%+ on Pixel 9 for tasks like AI Notification Summaries and Proofread; up to 55% better token acceptance on structured tasks.
  • Output is bit-for-bit identical to the base model -- incorrect drafts are discarded, so safety alignment is preserved.
  • Parallel work: Gemma 4 MTP drafters (trained jointly from scratch) deliver up to 3x speedup and are available on HuggingFace under Apache 2.0.
  • Key implication: MTP can now be added post-hoc to any deployed model, no pre-training required.

Getting a language model to run fast on a phone is already hard. Getting it to run fast without retraining it is a different problem entirely. That is exactly what Google Research's new work on frozen Multi-Token Prediction solves, and it is already live on Pixel 9 and 10 devices powering features like AI Notification Summaries and Proofread.

The bottleneck no one talks about

Standard language models generate text autoregressively, meaning they process and output just one word (or token) at a time. This step-by-step process creates a bottleneck, underutilizing the phone's processing power while straining its memory bandwidth, which can ultimately slow down the user experience and drain the battery.

The standard fix for this is speculative decoding -- a technique where a small, fast "drafter" model guesses several tokens ahead, and the main model verifies them all in one parallel pass. If the guesses are right, you get multiple tokens for the cost of one forward pass. Building on prior approaches like the EAGLE framework and Confident Adaptive Language Modeling (CALM), Google designed new architectural components to maximize these efficiency gains specifically for mobile environments.

But traditional speculative decoding has a hidden cost on mobile: running a separate "standalone" drafter model (e.g., 128M parameters) competes for limited RAM. Furthermore, a standalone drafter is "blind" to the main model's rich internal state, predicting next tokens based solely on text history without the semantic context the main model has already computed.

Grafting a head onto a frozen model

MTP addresses these inefficiencies by moving from a standalone architecture to an integrated one. Instead of training a separate small language model to draft tokens, a lightweight Transformer head -- the MTP head -- is appended to the final layers of the main model. This head takes the backbone's final internal representations and uses them to autoregressively predict a sequence of future tokens.

Architecture diagram showing frozen Gemini Nano backbone connected to a trainable MTP head with parallel token verification

The key constraint here is that this work focuses on retrofitting the drafter head to operate independently of the pre-training pipeline. Google takes a fully trained Gemini Nano v3 model, freezes its weights, and attaches a dense transformer stack -- the MTP head -- to the final layers, training only these parameters to minimize the prediction error on future tokens. The base model is never touched.

This "frozen backbone" approach has a critical safety property: because incorrect drafts are discarded during verification, the final output remains bit-for-bit identical to the main model, allowing efficiency updates to roll out with full backward compatibility. You get speed without any risk of behavior change.

The zero-copy trick

Memory is the real enemy on mobile. Even if you share weights between the drafter and the main model, a naive implementation still forces the drafter to build its own KV cache -- a running memory of all past context. On a phone, that is a serious problem.

To solve this, Google engineered a zero-copy architecture where the MTP head cross-attends directly to the main model's frozen KV cache. This allows the drafter to query the "memories" and context already computed by the backbone without duplication, eliminating drafter prefill latency and reducing the runtime memory footprint.

The concrete savings: 130MB per instance compared to a standalone drafter, by saving drafter embedding lookup tables, prefill dot attention variants, and application-specific tuning parameters. On a device with 8-12GB of total RAM shared across the OS, apps, and the model, 130MB is not trivial.

Where it shines -- and where it doesn't

MTP drafters consistently produce more accurate token predictions, resulting in speedups on Pixel 9 devices of 50% or more, depending on the task, compared to standalone drafters of comparable parameter counts. But the gains are not uniform across task types:

  • Instruction-following tasks (summarization, rewriting with complex constraints): MTP significantly outperformed standalone fine-tuned drafters, because it has access to the backbone's semantic understanding of the instruction.
  • Structurally predictable tasks (smart replies, templated outputs): the MTP head effectively learned the syntactic patterns of the main model, achieving up to a 55% improvement in token acceptance.
  • Open-ended generation: gains are more modest. When the output is highly uncertain and branching, the drafter's predictions are harder to accept, and the verification overhead eats into the speedup.

What it means in production

In production workloads such as AI Notification Summaries and Proofread, MTP correctly predicts an average of nearly two additional tokens per inference pass. Furthermore, fewer verification steps mean less time waking heavy processors, reducing energy consumption and improving battery life.

This is already deployed. Recently rolled out to the Pixel 9 and 10 series, this approach acts as an out-of-the-box speedup. For developers, it eliminates a major friction point: delivering high-speed on-device AI without the need to fine-tune separate, memory-heavy drafting models for every new task.

The bigger picture

This work sits within a broader Google push on MTP. MTP drafters for the Gemma 4 family deliver up to a 3x speedup without any degradation in output quality or reasoning logic -- but those were trained jointly with the backbone from the start. The Pixel work is harder: it proves you can bolt MTP onto a model that was never designed for it, without retraining, and still get substantial gains.

The assumption that has to update across the field: you no longer need to bake speculative decoding into a model at pre-training time to benefit from it. Frozen MTP retrofitting opens a path for any team running a deployed production model to add a lightweight drafter head post-hoc, without touching the base weights, without risking safety alignment, and without the memory overhead of a standalone drafter. That is a meaningful shift for anyone building on-device AI pipelines.

Google is looking forward to integrating MTP on future Pixel devices, as well as exploring alternative architectures -- including parallel decoding and paradigms without auxiliary heads -- to further drive down draft latency and increase simultaneous token verification under strict mobile constraints. Verification leniency -- relaxing the strict exact token match for specific use cases -- is also on the roadmap as a way to squeeze out further gains at the edge.

Trending
  • No trending articles

Comments

avatar

Next Reads