Models

Google DeepMind releases EmbeddingGemma 2 for phones

Google DeepMind has launched EmbeddingGemma 2, a modular 740-million-parameter model that enables fast, private multimodal search across text, audio, and video directly on consumer devices.

AlphaSignal1 day agoModels
Image: AlphaSignal

Google DeepMind has introduced EmbeddingGemma 2, a natively multimodal embedding model designed to run locally on consumer hardware under the Apache 2.0 license. The model features a modular architecture totaling 740 million parameters, which includes a 270-million-parameter text core, a 170-million-parameter vision encoder, and a 300-million-parameter audio encoder. This design allows developers to load only the text component or add the other encoders as needed. On a Pixel 11 Pro, the quantized text-only configuration requires about 191 megabytes of RAM, while the full multimodal setup uses approximately 567 megabytes.

In terms of processing capacity, the model can handle up to 5.5 minutes of audio, 29 images, or 58 video frames in a single pass. It utilizes Matryoshka Representation Learning, allowing developers to truncate its standard 768-dimensional output vectors down to 512, 256, or 128 dimensions at inference time to save memory. For example, storing one million raw 768-dimensional vectors at 32-bit precision takes up about 3.1 gigabytes, but reducing them to 128 dimensions shrinks that footprint to roughly 512 megabytes before database overhead.

On the Massive Text Embedding Benchmark, or MTEB, the model raised its MTEB Code score from 68.76 to 78.68, a jump of 9.92 points. This improvement makes it highly competitive for local code search, outperforming some specialist models more than twice its size. Furthermore, because EmbeddingGemma 2 shares its text tokenizer and audio encoder with Gemma 4, compatible runtimes can reuse these components to lower the overall memory footprint of on-device retrieval-augmented generation pipelines.

The model weights are currently available on Hugging Face and Kaggle, with support for runtimes like Ollama, vLLM, MLX, Transformers, Sentence Transformers, llama.cpp, SGLang, and LM Studio. For edge deployment, developers can use MediaPipe or LiteRT. Google has showcased these capabilities in reference applications such as Instant Media Search, Video Moments Finder, and Google AI Edge Foresight, which pairs the new model with Gemma 4.

This is our own summary of reporting by AlphaSignal

More in Models