Google DeepMind's EmbeddingGemma 2 Searches Audio, Video, and Images on Your Phone
Google DeepMind's new 740M-parameter open model unifies text, code, images, audio, and video embeddings in a single on-device package.
- Google DeepMind released EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model for on-device use.
- Handles text, code, images, audio, and video in one shared embedding space, under Apache 2.0.
- Modular: 270M text core plus optional 170M vision and 300M audio encoders.
- Runs in ~191MB RAM text-only or ~567MB full multimodal on a Pixel 11 Pro with quantization.
- MTEB Code jumps 9.92 points to 78.68; outperforms some models over 2x its size.
- Weights live on Hugging Face and Kaggle, with Ollama, vLLM, and MLX support.
EmbeddingGemma 2 brings multimodal retrieval to consumer devices
Google DeepMind has released EmbeddingGemma 2, a 740 million-parameter embedding model that processes text, images, audio, and video on consumer hardware. Google describes it as its first natively multimodal embedding model, designed for cross-modal queries such as using a voice memo to search videos or text to retrieve an audio clip.
Embedding models convert content into numerical vectors, placing semantically related items near one another. Running that process locally can reduce network latency, support offline search, and keep personal data on the device.
The text-only EmbeddingGemma 1 passed 20 million downloads. Its successor adds three optional modality encoders while remaining below one billion parameters, a practical threshold for phones and laptops with limited memory.
740M parameters, loaded by modality
| Specification | EmbeddingGemma 2 |
|---|---|
| Architecture | Gemma 4 |
| License | Apache 2.0 |
| Parameters | 740M total: 270M for text, 170M for vision, and 300M for audio |
| Vector sizes | 768, 512, 256, or 128 dimensions |
| Context window | 8K tokens, four times EmbeddingGemma 1 |
| Approximate capacity | 5.5 minutes of audio, 29 images, or 58 video frames per pass |
| Reported Pixel memory | About 191MB for quantized text weights and 567MB for the full multimodal model on a Pixel 11 Pro |
The modular design lets applications load the 270M text component alone or add the vision and audio encoders as needed. Actual application memory will exceed the reported model figures once the runtime, activations, vector index, and input data are included.
Vectors that shrink after training
Matryoshka Representation Learning trains the beginning of each output vector to remain useful on its own. Developers can truncate a 768-dimensional embedding to 512, 256, or 128 dimensions at inference time, trading some retrieval accuracy for lower storage and memory use without retraining the model.
At 32-bit precision, one million raw 768-dimensional vectors occupy about 3.1GB. Reducing them to 128 dimensions cuts that figure to roughly 512MB before vector-database overhead, a sixfold reduction.
Code retrieval posts the clearest gain
Google reports that EmbeddingGemma 2 retains its predecessor’s multilingual text performance while raising its MTEB Code score from 68.76 to 78.68. MTEB, the Massive Text Embedding Benchmark, evaluates how well embedding models support tasks such as retrieval, clustering, and classification.
| MTEB Code result | Score |
|---|---|
| EmbeddingGemma 1 | 68.76 |
| EmbeddingGemma 2 | 78.68 |
| Improvement | 9.92 points |
That gain makes the model relevant to local code search and retrieval for coding agents. Google also reports leading quality per parameter among sub-1B models across image, video, document, and audio evaluations, including results above some specialist models with more than twice as many parameters.
Benchmark scores depend on datasets, preprocessing, vector dimensions, and quantization settings, so production evaluations should use representative data and target hardware. Larger embedding models can still deliver higher accuracy, and the full multimodal configuration consumes more than half a gigabyte of reported active RAM before application overhead.
One stack for retrieval and generation
Retrieval-augmented generation, or RAG, first searches an embedding index for relevant material and then supplies those results to a generative model. Because EmbeddingGemma 2 shares Gemma 4’s text tokenizer and audio encoder, compatible runtimes can reuse those components across retrieval and generation.
Component reuse can lower the combined memory footprint of a phone-resident RAG pipeline, especially for applications that search audio and then generate an answer. The savings depend on whether the selected runtime shares model components instead of loading separate copies.
Google has published reference applications in the Google AI Edge Gallery. Instant Media Search retrieves library items from a text or image query, while Video Moments Finder locates passages in video from text or audio. The Google AI Edge Foresight app demonstrates EmbeddingGemma 2 paired with Gemma 4 for contextual reasoning.
Routes to production
The model weights are available from Hugging Face and Kaggle. Google says support for the Gemini Enterprise Agent Platform Model Garden will follow.
- General inference: Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LM Studio
- Edge deployment: MediaPipe for packaged retrieval tasks and LiteRT for custom integrations
- Browser testing: A hosted WebGPU demo
- Fine-tuning: A published Unsloth guide
Where local retrieval pays off
Applications that handle private, offline, or latency-sensitive data can use EmbeddingGemma 2 to keep retrieval close to the source. Suitable workloads include:
- Semantic search across notes, photos, documents, and voice memos stored on a device
- Local codebase retrieval for coding agents running on a laptop
- Cross-modal search, such as finding the moment in a video when someone says a specific phrase
- On-device classification and routing through the MediaPipe Decision Task API
Managed cloud retrieval remains a reasonable fit when data already resides on a server and network latency, offline access, and device privacy are secondary concerns. Mobile and desktop applications that search user-owned content gain more from the shared Gemma 4 components, adjustable vector dimensions, and reported sub-600MB multimodal memory footprint.