Google Launches EmbeddingGemma 2 with 740M Parameters and Open License
- Published
- Oct 6, 2026 — 19:57 UTC
EmbeddingGemma 2, launched on October 6, 2026, features 740 million parameters and is designed for on-device multimodal embeddings. Developed by Sahil Dua and Henrique Schechter Vera at Google DeepMind, this model follows the original EmbeddingGemma, which saw over 20 million downloads since its release last year. EmbeddingGemma 2 achieves a maximum storage reduction of 6x for local vector databases using MRL, enhancing efficiency for developers.
The model supports an 8K token context window, allowing for the processing of up to 5.5 minutes of audio, 29 images, and 58 video frames simultaneously. On a Google Pixel 11 Pro, it requires 191MB of active RAM for text-only weights and 567MB for the full multimodal model. Performance benchmarks indicate a 9.92-point improvement in MTEB Code scores, rising from 68.76 to 78.68.
EmbeddingGemma 2 is available under the Apache 2.0 license, facilitating broader accessibility and integration into various applications. It can be accessed through platforms like Hugging Face and Kaggle for model weights, while the Google AI Edge Gallery serves as a testing platform. The MediaPipe Decision Task API is included for classification and routing tasks, and LiteRT supports on-device search and retrieval-augmented generation systems. This release positions EmbeddingGemma 2 as a leading choice for developers focused on local data privacy and reduced latency.
By Turing Wire Newsdesk · Oct 6, 2026 · How we work →
Summarised from Google DeepMind Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.
Source: Google DeepMind Blog
