Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio
Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.
Google DeepMind has released EmbeddingGemma 2, an open multimodal embedding model with 740M total parameters that maps text (including code), images, video and audio, alone or combined, into one shared 768-dimensional vector space. The weights are available on Hugging Face under the Apache 2.0 license.
Key specifications
- Parameters: 740M total. The 270M text model (130M transformer backbone + 140M embedder) is paired with a 170M vision encoder and a 300M audio encoder.
- Context window: 8,192 tokens. Google says this covers minutes of audio or video.
- Output dimension: 768 native, with Matryoshka Representation Learning (MRL) truncation to 512, 256 and 128 dimensions.
- Architecture: 24 layers, model dimension 512, hidden dimension 2048, 4 attention heads, GQA/MQA attention, 1024-token sliding window with a 5:1 local-to-global ratio, 262,144-token vocabulary, mean pooling and a 512→768 projection layer.
- Languages: 100+.
- Foundation: Built on the architecture and capability advances of Gemma 4, according to Google.
- Pricing: Not applicable for the open weights. Hosted-API pricing has not been disclosed.
- Training cutoff: Not disclosed in the source material.
Benchmarks (Google-reported, full-precision, 768d)
| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|
| MTEB multilingual v2 (mean task) | 61.36 | 61.15 |
| MTEB code v1 (NDCG@10) | 78.68 | 68.76 |
| MTEB English v2 (mean task) | 68.46 | not listed |
| MIEB lite | 64.64 | n/a |
| MMEB v2 Image (Hit@1) | 57.28 | n/a |
| MMEB v2 VisDoc (NDCG@5) | 67.84 | n/a |
| MMEB v2 Video (Hit@1) | 50.67 | n/a |
| MMEB v2 Overall | 59.01 | n/a |
| MSEB Retrieval (MRR@10) | 69.54 | n/a |
| MAEB | 49.39 | n/a |
The code score is a gain of roughly 14% over the predecessor, matching Google's claim. The multilingual text gain is 0.21 points. The model card includes no comparisons against third-party embedding models.
Truncation trade-offs
Truncated vectors must be re-normalized, and queries and documents must share a dimension. Google's figures show:
- 512d (1:1.5 compression): MTEB multilingual 61.17, MMEB v2 overall 58.38.
- 256d (1:3): MTEB multilingual 60.41, MMEB v2 overall 56.24.
- 128d (1:6): MTEB multilingual 57.89, MMEB v2 overall 45.65.
Google recommends 128d only for text-only workloads.
Selective encoder loading
The vision and audio encoders load independently through SentenceTransformer's config_kwargs:
- Text only: 270M
- Text and image: 440M
- Text and audio: 570M
- Full multimodal: 740M
The model uses short task-instruction prefixes for text inputs, such as task: search result | query: {query}, covering search, question answering, fact checking, code retrieval, classification, clustering and sentence similarity. Prefixes apply to text only. Omitting them still works but reduces precision, according to the model card.
What this means
The main change is scope. EmbeddingGemma 2 puts four modalities into one sub-1B model with a shared index, which lets developers build cross-modal retrieval, such as text queries over video or audio, without stitching together separate models. The modular encoders keep the text-only footprint at 270M parameters, so existing on-device text deployments need not pay for modalities they don't use.
For text-only users, the upgrade case rests on code retrieval. The 10-point MTEB code gain is substantial, but the 0.21-point multilingual gain is marginal. Embeddings from different models are not interchangeable, so any move from EmbeddingGemma 1 means re-indexing the corpus.
Treat the numbers as vendor-reported until independent evaluations appear. The card offers no head-to-head results against other open multimodal embedders, and the video (50.67) and audio (MAEB 49.39) scores suggest those modalities are less mature than text and image. The sharp MMEB drop at 128d means aggressive truncation is a text-only optimization.
Related Articles
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says it runs in ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model on a quantized Pixel 11 Pro. Google also released a Mac app, AI Edge Foresight, to demonstrate it.
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.
Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context
Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.
Mistral releases Large 4, a 1-trillion-parameter multimodal model, with open weights due in three weeks
Mistral AI released Mistral Large 4 (ML4), a multimodal model with one trillion parameters, on Tuesday. It is currently available only through a public guardrail endpoint, and Mistral plans to publish the weights in about three weeks after safety testing. Benchmark results, pricing and context window have not been disclosed.
Comments
Loading...