Google's EmbeddingGemma 2 puts text, code, images and audio in one open 740M-parameter embedding model
The biggest measured gain is in code retrieval. Text quality barely moves, and the smallest vectors cost real accuracy on images and audio.
Google DeepMind released EmbeddingGemma 2 on October 6, an open embedding model that maps text, code, images, video and audio into a single 768-dimensional vector space. It has 740 million parameters, is built on the Gemma 4 architecture and ships under Apache 2.0. Weights are on Hugging Face and Kaggle, according to Google's announcement; availability in Google's Model Garden is listed as coming soon.
This is not a chat model. Embedding models turn content into vectors so software can find similar things: the retrieval step in a RAG pipeline, semantic search over a codebase, or routing a request to the right tool. That part of the stack rarely gets attention, but it decides what a language model actually gets to read.
What is in the release
- Modular size. The text backbone is 270M parameters (130M transformer plus 140M embedder). The vision encoder (170M) and audio encoder (300M) are optional and can be left unloaded, which gives four configurations: 270M text-only, 440M text plus images, 570M text plus audio, 740M for everything.
- 8,192-token context, four times the first EmbeddingGemma. All modalities share that budget at fixed rates, per the model card: 280 tokens per image, 140 per video frame, 25 per second of audio. Google puts the ceiling at about 5.5 minutes of audio, 29 images or 58 video frames.
- Matryoshka truncation. Vectors can be cut from 768 to 512, 256 or 128 dimensions, for up to 6x less vector storage.
- On-device memory. Google says that with quantization on a Pixel 11 Pro it needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.
- Tooling on day one: sentence-transformers, transformers, vLLM, SGLang, llama.cpp, Ollama, LM Studio, MLX, transformers.js, LiteRT and MediaPipe, with fine-tuning guidance from Unsloth.
The numbers, and what they do not show
All benchmark figures come from Google's own testing, reported in the model card on the full-precision checkpoint. No independent evaluation was available at publication time.
The headline gain is code. On MTEB Code, EmbeddingGemma 2 scores 78.68 against 68.76 for its predecessor, a 9.92-point increase that Google describes as about 14%. On multilingual text, the change is small: 61.36 on MTEB multilingual v2, against 61.15 for the first version. In practice, text retrieval quality is roughly where it was; the new value is code, plus image, video and audio support in the same model.
For the new modalities there is no predecessor to compare with. The card lists 64.64 on MIEB lite (images), 50.67 on MMEB v2 video, 69.54 on MSEB audio retrieval and 49.39 on MAEB. Google calls these leading scores among multimodal embedders under 1B parameters and says the model beats some specialist models more than twice its size, but the announcement does not name those rivals.
Truncation has a cost, and the card says so directly. At 256 dimensions, the losses are modest: multilingual text drops from 61.36 to 60.41 and code from 78.68 to 76.18. At 128 dimensions, multimodal quality falls much more: the overall MMEB v2 score goes from 59.01 to 45.65, and audio retrieval from 69.54 to 56.71. Google recommends 128 dimensions mainly for text-only workloads, tested against your own data.
What changes for developers
Three practical points stand out for anyone building retrieval or coding agents.
Task prefixes matter. Like its predecessor, the model is trained with short instructions in front of text inputs, for example task: code retrieval | query: ... for a code-search query and title: <filename> | text: <code> for the indexed file. The card warns that leaving the prefix out still works but reduces precision. Images, audio and video take no prefix. If you are swapping in a new embedding model and see worse results, check this first.
Local code retrieval gets cheaper. A 270M text-and-code embedder that scores 78.68 on MTEB Code is small enough to index a repository on a laptop. That makes it a reasonable candidate for coding-agent setups that keep source code off third-party APIs. Whether it beats the hosted embedding APIs you already use is a question for your own evaluation set, not the vendor table.
Precision and normalization traps. The card says float16 can produce NaN or degraded outputs; use bfloat16 or float32. After truncating, vectors must be re-normalized, and queries and documents must use the same dimension. Skipping normalization does not throw an error; it just gives plausible-looking but worse rankings. According to the developer guide, one million 768-dimensional vectors take about 1.5GB, against about 250MB at 128 dimensions.
The model also shares its text tokenizer and audio encoder with Gemma 4, so Google says the two can run together in one on-device pipeline with a lower combined memory footprint. That claim matters mostly for mobile and edge apps that want local RAG without a server.
Why this matters
- Embedding models decide what a RAG system or coding agent retrieves, so a stronger small code embedder directly affects local codebase search.
- An Apache 2.0 model that runs text-only in about 191MB of RAM (Google's figure) makes fully local retrieval practical for privacy-sensitive apps.
- One vector space for text, images, audio and video lets a single index handle cross-modal search without stitching together separate models.
Key takeaways
- 740M parameters total, with a 270M text-only configuration; vision and audio encoders are optional.
- Google reports MTEB Code at 78.68, up from 68.76; multilingual text moves only from 61.15 to 61.36.
- 8,192-token context shared across modalities; vectors truncate to 512, 256 or 128 dimensions, with large multimodal losses at 128.
- Use the task prefixes and bfloat16 or float32, and re-normalize after truncation, or quality drops without warning.
Sources
- GooglePrimaryEmbeddingGemma 2: an open, lightweight multimodal embedding modelblog.google
- Google DeepMind (Hugging Face)Primarygoogle/embeddinggemma-2 model cardhuggingface.co
- Google Developers BlogPrimaryEmbeddingGemma 2: The Developer Guidedevelopers.googleblog.com
- Google AI for DevelopersPrimaryEmbeddingGemma 2 model cardai.google.dev
- embeddings
- rag
- open-weights
- on-device
- code-search
- multimodal
- Google DeepMind
- EmbeddingGemma 2
- EmbeddingGemma
- Gemma 4