Google releases EmbeddingGemma 2, a 740M-parameter open multimodal embedding model scoring 78.68 on MTEB Code
Google released EmbeddingGemma 2, an open 740-million-parameter model that embeds text, images, video, audio, and code. It scores 78.68 on MTEB (Code), up from 68.76 for its predecessor, and Google claims it beats models up to twice its size on multimodal embedding benchmarks.
Google has released EmbeddingGemma 2, an open-weights embedding model with 740 million parameters that maps text, images, video, audio, and code into numerical vectors. Google claims it is the most compact model of its kind and that it outperforms competing models up to twice its size on multimodal embedding benchmarks.
Key specifications
- Parameters: 740 million (multimodal version); a 270-million-parameter version is available for text-only tasks
- Modalities: text, images, video, audio, code
- MTEB (Code) score: 78.68, versus 68.76 for the predecessor, a gain of nearly 10 points
- Memory footprint: around 191 MB of RAM, according to Google
- Browser latency: roughly 20 to 70 milliseconds per query via WebGPU, according to Google
- Storage: Google says local vector database storage drops by up to six times
- Availability: weights on Hugging Face and Kaggle, with a developer guide and documentation
- Pricing: no API pricing; the model is distributed as open weights and runs locally without an API key
- Context window, embedding dimensions, license terms, and training cutoff: not disclosed in the available information
Benchmark results
On the Massive Text Embedding Benchmark (Code), EmbeddingGemma 2 scores 78.68. Google says that puts it on par with much larger models. The comparison models, their parameter counts, and scores on the multimodal benchmarks behind the "twice its size" claim were not itemized. The claim comes from Google and has not been independently verified.
Local and offline deployment
The model is built to run on-device. Google says a query takes about 20 to 70 milliseconds in the browser through WebGPU, with a memory requirement of about 191 MB. The company pairs it with small open models such as Gemma 4 to build retrieval-augmented generation (RAG) apps that run offline, with no data sent to external servers.
The 270-million-parameter variant covers text-only workloads, where the multimodal capacity of the larger model is not needed.
What this means
Embedding models are a quiet bottleneck in RAG pipelines. They decide what gets retrieved, and most strong ones sit behind APIs or need server-class hardware. A model of 740M parameters, or 270M for text, with a roughly 191 MB runtime footprint moves retrieval into the browser and onto edge devices, which matters for privacy-sensitive and offline applications.
The nearly 10-point gain on MTEB (Code) over the predecessor is large for a single generation. Other labs' models will still be the reference for pure text retrieval, and the headline claim of beating models twice its size rests on Google's own comparisons. Developers should test it on their own data, especially for cross-modal retrieval, before replacing an existing embedding stack.
The open weights and the single-vector-space approach to mixed media also lower the cost of building search across documents, screenshots, audio, and video without a separate model per modality. Missing details, notably license terms, context length, and embedding dimensionality, will determine how easily teams can adopt it commercially.
Related Articles
Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio
Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says it runs in ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model on a quantized Pixel 11 Pro. Google also released a Mac app, AI Edge Foresight, to demonstrate it.
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.
Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context
Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.
Comments
Loading...