Google launches EmbeddingGemma 2, an open 740M-parameter multimodal embedder

Google has launched EmbeddingGemma 2, an open embedding model that goes beyond text to unify code, images, video and audio in a shared embedding space. It is built on the Gemma 4 architecture, has 740 million parameters, and is released under a commercially permissive Apache 2.0 license. Google calls that size optimal for on-device inference.
The announcement starts from the first EmbeddingGemma, introduced last year as a lightweight option for text embeddings on consumer hardware. Google says the response beat its expectations: more than 20 million downloads, with builders using it for on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines. Google describes the new model as built from the same technology as its Gemini Embedding models. Its pitch is that a single, natively multimodal model can find a specific video clip from a voice memo, or search hours of audio recordings from a text query.
Google lists five properties. First, it claims leading scores among sub-1B multimodal embedders for its size on benchmarks such as MTEB Code and MAEB (the Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision and audio tasks. Second, the model is modular: text-only workloads need as little as 270M parameters, with optional vision (170M) and audio (300M) encoders for full multimodal support. Third, Matryoshka Representation Learning (MRL) lets developers truncate output vectors from 768 dimensions down to 512, 256 or 128, for up to 6x storage reduction in local vector databases and memory usage. Fourth, with quantization on a Google Pixel 11 Pro, it needs as little as about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model. Fifth, it has an 8K token context window, 4x larger than EmbeddingGemma 1, enough for up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.
On quality, Google says EmbeddingGemma 2 matches the multilingual text performance of the first model while improving MTEB Code by 9.92 points, from 68.76 to 78.68. It says that makes the model suited to local codebase indexing, semantic code search and coding agent retrieval. Across image, video, document and audio tasks, Google says it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size. Full evaluation metrics are in the model card.
Google's argument for running this locally is that generating embeddings on the device helps ensure data privacy, reduces pipeline latency, and lets developers build cross-modal search and retrieval that works entirely offline. Paired with a generative model such as Gemma 4, it enables on-device RAG over multimodal data. Because EmbeddingGemma 2 shares Gemma 4's text tokenizer and audio encoder, the two can run in one pipeline with a lower combined memory footprint.
On availability, the weights are on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. The LiteRT Community on Hugging Face hosts versions optimized for on-device use. Deployment options include Google AI Edge MediaPipe, LiteRT, transformers.js and WebGPU. Serving works with transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio. Qdrant is named for storing vectors, and Unsloth provides fine-tuning guidance.
Key facts
- EmbeddingGemma 2 has 740 million parameters, is built on Gemma 4, and unifies code, images, video and audio in one embedding space under an Apache 2.0 license.
- It is modular: as little as 270M parameters for text-only use, with optional 170M vision and 300M audio encoders.
- Google reports MTEB Code rising 9.92 points over EmbeddingGemma, from 68.76 to 78.68, and says it outperforms some specialist models more than twice its size.
- MRL truncation takes vectors from 768 down to 512, 256 or 128 dimensions for up to 6x storage reduction; quantized, it needs about 191MB (text-only) or 567MB (full multimodal) of RAM on a Pixel 11 Pro.
- The 8K token context window is 4x larger than the first model's; weights are on Hugging Face and Kaggle, and Model Garden availability is coming soon.
Why it matters
The first EmbeddingGemma covered text only. This release adds code, images, video and audio to the same embedding space, and Google says it still fits on consumer hardware. That means one small model can handle cross-modal search, such as finding a video clip from a voice memo, without a server. Google also reports more than 20 million downloads of the first model, which suggests the audience for small local embedders is large. The Apache 2.0 license allows commercial use.
Who it affects
Developers building on-device search and RAG pipelines are the main audience, along with teams that want embeddings generated locally for data privacy or offline use. Google also points to people building local codebase indexing, semantic code search and coding agent retrieval, given the MTEB Code gain. Teams already running Gemma 4 may benefit from the shared text tokenizer and audio encoder, which Google says lowers combined memory use.
How to use it
The weights are on Hugging Face and Kaggle; Gemini Enterprise Agent Platform Model Garden availability is coming soon. The LiteRT Community on Hugging Face has models optimized for on-device use. For apps, Google points to Google AI Edge MediaPipe for turnkey embedding and retrieval, LiteRT for custom integration, and transformers.js or WebGPU for the browser. For serving, it lists transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio. Qdrant is named for vector storage and Unsloth for fine-tuning guidance. To cut storage, truncate vectors from 768 to 512, 256 or 128 dimensions with MRL. The license is Apache 2.0, and no price is given.
How solid is it
This is a first-party launch post, and the performance claims are Google's own. The only benchmark figures quoted are the MTEB Code scores (68.76 to 78.68); the MAEB and the vision and audio results are not quoted, and the post points to the model card for full metrics. No named competitor models are compared against. The RAM figures of about 191MB and 567MB are for quantized models on a Pixel 11 Pro, and no latency or throughput figures are given.
Risks and caveats
Phrases like best-in-class for its size and leading scores among sub-1B multimodal embedders are Google's characterisation, and the post does not name the models behind them. The post does not say how the 740M total relates to the 270M text-only figure and the 170M and 300M encoders, so the footprint of a given configuration is best checked in the model card. Model Garden availability has no date beyond coming soon. The post gives no calendar date for the launch.
“Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.”
— Google, EmbeddingGemma 2 announcement