Unsloth publishes GGUF quantizations of Google DeepMind's 740M-parameter EmbeddingGemma 2 multimodal embedding model
Unsloth has released GGUF quantizations of EmbeddingGemma 2, Google DeepMind's open 740M-parameter embedding model for text, images, video and audio. The base model maps all four modalities into one 768-dimensional space with an 8K context window under Apache 2.0.
Unsloth has published GGUF quantizations of EmbeddingGemma 2, Google DeepMind's open-weight embedding model that maps text, images, video and audio into a single 768-dimensional vector space. The upload is a community repackaging for local runtimes. The underlying model is by Google DeepMind and is licensed under Apache 2.0.
What Unsloth released
The repository, unsloth/embeddinggemma-2-GGUF, holds quantized versions of Google's checkpoint. Unsloth claims its Dynamic 3.0 quantization method "achieves superior accuracy & outperforms other leading quants." That is a company claim, and the source gives no quantization-specific benchmarks. Quant levels and file sizes are not stated in the source material.
All benchmark figures below come from Google's full-precision checkpoint, not the GGUF builds.
EmbeddingGemma 2 specifications
- Parameters: 740M total. This is a 270M text model (130M transformer plus 140M embedder), a 170M vision encoder and a 300M audio encoder.
- Architecture: 24 layers, model dimension 512, hidden dimension 2048, 4 attention heads, GQA/MQA attention, 1024-token sliding window, 5:1 local-to-global ratio, 262,144-token vocabulary, mean pooling, and a 512→768 projection layer.
- Context window: 8,192 tokens. Google says this is enough for minutes of audio or video.
- Output: 768 dimensions natively. Matryoshka Representation Learning (MRL) supports truncation to 512, 256 and 128 dimensions.
- Languages: 100+. Google also says code tasks improve about 14% over the predecessor.
- Base: Built on Gemma 4 architecture advances, per Google.
- Pricing: Not applicable for local weights. No hosted API pricing was disclosed. The training cutoff date was not disclosed.
The vision and audio encoders load selectively. Effective size is 270M for text only, 440M for text and image, 570M for text and audio, and 740M for full multimodal.
Reported benchmarks (768d, full precision)
| Benchmark | EmbeddingGemma 2 | EmbeddingGemma 1 |
|---|---|---|
| MTEB multilingual v2 | 61.36 | 61.15 |
| MTEB code v1 (NDCG@10) | 78.68 | 68.76 |
| MIEB lite | 64.64 | n/a |
| MMEB v2 Image (Hit@1) | 57.28 | n/a |
| MMEB v2 VisDoc (NDCG@5) | 67.84 | n/a |
| MMEB v2 Video (Hit@1) | 50.67 | n/a |
| MSEB Retrieval (MRR@10) | 69.54 | n/a |
| MAEB | 49.39 | n/a |
Text quality is nearly flat against the first-generation model on multilingual MTEB (+0.21). The gain is on code retrieval (+9.92 points). The image, video and audio benchmarks have no first-generation baseline because the original model was text-only.
Truncation trade-offs
Google's truncation table shows how quality degrades as vectors shrink:
- 512d: MTEB multilingual 61.17, MMEB v2 overall 58.38
- 256d: 60.41 and 56.24
- 128d (1:6 compression): 57.89 and 45.65
Google says quality impact is minimal down to 256d and that 128d suits text-only workloads. Multimodal quality drops sharply at 128d: MMEB v2 falls from 59.01 to 45.65, and MSEB retrieval from 69.54 to 56.71.
Usage requirements
The model uses task instruction prefixes on text inputs only, such as task: search result | query: {query} for search queries and title: {title} | text: {content} for documents. Omitting prefixes still works but reduces precision, per Google. Images, video and audio take no prefix. Truncated vectors must be re-normalized (L2) before cosine similarity, or ranking degrades silently. Selective encoder loading differs across libraries, so GGUF runtime behavior should be checked against each tool's documentation.
What this means
A single embedding space for four modalities in under 1B parameters, with an Apache 2.0 license, lowers the barrier to on-device multimodal retrieval. The selectable 270M text-only footprint keeps it practical for phones and laptops. The GGUF release matters mainly for the llama.cpp ecosystem and similar local runtimes, though multimodal support in those tools may lag text-only support. That has not been verified here.
The open questions are quantization loss and tooling support. Embedding models are sensitive to precision because small vector shifts change nearest-neighbor rankings. Until Unsloth or third parties publish MTEB or MMEB results for the quantized files, treat Google's numbers as an upper bound. Teams should also measure retrieval quality on their own data at the quant level they plan to deploy.
Related Articles
Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio
Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.
Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
DeepMind essay argues AGI will emerge from human-agent networks, not a lone superintelligence
Google-affiliated researchers Benjamin Bratton, Blaise Agüera y Arcas and James Manyika propose "Artificial Symbiotic Intelligence," a framework in which AGI emerges from a social system of people and AI agents rather than a single self-improving machine. The essay, written for the Deepmind Institute, is a conceptual argument and reports no benchmark results.
Google's SynthID detector goes global in English; Apple Intelligence watermarking to follow 'soon'
Google has moved its SynthID AI content detector from limited early access to global availability in English. The tool now covers content from OpenAI, Nvidia, Kakao and, "soon," Apple Intelligence, in addition to Google's own Gemini models.
Comments
Loading...