Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says it runs in ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model on a quantized Pixel 11 Pro. Google also released a Mac app, AI Edge Foresight, to demonstrate it.
Google announced EmbeddingGemma 2 on October 6, 2026, a 740-million-parameter, natively multimodal embedding model built on the Gemma 4 architecture and released under an Apache 2.0 license. Google says the model is designed to "organize, search, and connect information directly on consumer hardware."
Key specifications
- Parameters: 740 million
- Architecture: Gemma 4
- License: Apache 2.0
- Modalities: Text, audio and video, handled by a single natively multimodal model, according to Google
- Memory footprint (quantized, Pixel 11 Pro): ~191MB active RAM for text-only weights; ~567MB for the full multimodal model, per Google
- Context window: Not yet disclosed
- Embedding dimensions: Not yet disclosed
- Benchmark scores: Not yet disclosed
- Pricing: No hosted API pricing disclosed. Weights are available under Apache 2.0.
- Training cutoff: Not disclosed
Google claims the model can search "through hours of audio recordings based on a text query" and locate specific video clips, all in one model rather than separate per-modality pipelines. These capabilities are Google's description. No independent evaluations or published retrieval benchmarks accompanied the announcement.
On-device focus
The memory figures target tight resource constraints. Google's numbers are for a quantized build on a Google Pixel 11 Pro. Demos are available in the Google AI Edge Gallery app on Android and iOS.
AI Edge Foresight for Mac
To demonstrate the model, Google released Google AI Edge Foresight, a Mac notetaking app. According to Google:
- Shorthand expansion: The app listens during video or in-person meetings while the user jots shorthand bullets. It then uses EmbeddingGemma 2 to "enrich" them into polished notes using the meeting transcript.
- Local processing: Transcript and audio are processed entirely on-device, not in the cloud.
- Personal knowledge retrieval: Users can point the app at project folders, reference materials, diagrams and a calendar. Local files and Google Drive are supported. A chat feature offers "conversational retrieval and summary."
- Live Assistance: A mode that will "listen and automatically answer questions."
Foresight joins two other Google Mac apps in the same family: AI Edge Gallery and AI Edge Eloquent, an offline transcription app.
The announcement does not say which model generates the polished notes or chat answers. EmbeddingGemma 2 is an embedding model, so it handles retrieval and matching rather than text generation. Google has not specified the generative component in the material available.
What this means
EmbeddingGemma 2 puts a single multimodal embedding model under a permissive license at a size that fits in a few hundred megabytes of RAM on a phone. If Google's memory figures hold up, developers building local retrieval over voice memos, meetings and video no longer need separate text and audio embedding stacks or a cloud round trip. The Apache 2.0 license also removes the usage restrictions that complicate commercial deployment of some open-weight models.
The main gaps are evidence and specifics. Without published retrieval benchmarks, context length or embedding dimensions, developers can't yet compare it against existing open embedding models. The memory claims also come from one device, the Pixel 11 Pro, with quantization, so performance on older phones and laptops is untested. Foresight is the more telling signal. Google is positioning on-device retrieval as the foundation for private, offline productivity apps, a use case where cloud-only embedding APIs are weakest.
Related Articles
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.
Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio
Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.
Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context
Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.
Mistral releases Large 4, a 1-trillion-parameter multimodal model, with open weights due in three weeks
Mistral AI released Mistral Large 4 (ML4), a multimodal model with one trillion parameters, on Tuesday. It is currently available only through a public guardrail endpoint, and Mistral plans to publish the weights in about three weeks after safety testing. Benchmark results, pricing and context window have not been disclosed.
Comments
Loading...