EmbeddingGemma 2 maps text, code, images, audio and video into one shared embedding space, and Google DeepMind has released it with open weights under an Apache 2.0 license. It is sized for phones and laptops rather than data centers.
Google DeepMind describes the 740 million parameter EmbeddingGemma 2 as the successor to last year's text-only EmbeddingGemma, which the company says passed 20 million downloads. This version is built on the Gemma 4 architecture. Developers load only the parts they need: roughly 270 million parameters covers text alone, and optional vision (170 million) and audio (300 million) encoders add the other modalities, Google DeepMind says.
The context window grew to 8K tokens, four times the first model's. With quantization on a Pixel 11 Pro, Google reports about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model. The company says the window covers up to 5.5 minutes of audio, 29 images or 58 video frames, or a mix. Matryoshka Representation Learning lets developers shrink the 768-dimension output vectors to 512, 256 or 128 dimensions, which Google puts at up to six times less storage for a local vector database.
On quality, Google claims EmbeddingGemma 2 posts the best scores among multimodal embedders under one billion parameters on benchmarks that include MTEB Code and the audio-focused MAEB, and says it beats some specialist models more than twice its size. The one before-and-after figure it gives is MTEB Code, up from 68.76 to 78.68, while multilingual text performance matches the first release. These are Google's own evaluations, so independent testing will matter.
One design detail is worth noting for anyone building offline search. EmbeddingGemma 2 shares its text tokenizer and audio encoder with Gemma 4, so the two can run in a single pipeline with a smaller combined memory footprint than separate models would need. Weights are on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability listed as coming soon. Google names support in transformers, sentence-transformers, MLX, vLLM, llama.cpp, Ollama and LM Studio, plus Qdrant for storing vectors.
For a team indexing a code repository or a media library on local hardware, the appeal of EmbeddingGemma 2 is one embedder in place of three. The memory figures above come from a single phone, so they are the numbers to retest on the target device.













