Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.
Google announced EmbeddingGemma 2 on October 6, 2026, a 740-million-parameter, natively multimodal embedding model built on the Gemma 4 architecture and released under an Apache 2.0 license. Google says the model is designed to "organize, search, and connect information directly on consumer hardware."
Key specifications
- Parameters: 740 million
- Architecture: Gemma 4
- License: Apache 2.0
- Modalities: Text, audio and video, handled by a single model, according to Google
- Memory footprint (quantized, Pixel 11 Pro): ~191MB active RAM for text-only weights; ~567MB for the full multimodal model
- Context window: not yet disclosed
- Embedding dimensions: not yet disclosed
- Pricing: not applicable for the open-weight release; no hosted API pricing has been disclosed
- Benchmark scores: none published in the announcement
- Training cutoff: not yet disclosed
The memory figures come from Google and apply to a quantized model on Google's own hardware. Independent verification is not yet available.
Capabilities
Google claims EmbeddingGemma 2 can "find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model." That points to a shared embedding space across text and other media, which removes the need to chain separate transcription and text-embedding steps.
The model is aimed at on-device retrieval under "tight resource constraints." Demos are available in the Google AI Edge Gallery app on Android and iOS.
AI Edge Foresight for Mac
To show the model in use, Google released Google AI Edge Foresight, a Mac notetaking app. According to Google:
- The app listens in the background during video or in-person meetings while the user types shorthand bullets.
- It uses EmbeddingGemma 2 to turn that shorthand into "polished notes using the meeting transcript instantly."
- Transcript and audio are processed entirely on the device, not in the cloud.
- Users can point the app to project folders, reference materials, diagrams and calendar data. It supports local files and Google Drive.
- A chat feature offers "conversational retrieval and summary with your personal knowledge."
- A Live Assistance feature will "listen and automatically answer questions."
Foresight joins two other Google Mac apps in the AI Edge family: AI Edge Gallery and AI Edge Eloquent, an offline transcription app. The announcement does not specify which model handles chat responses or Live Assistance. EmbeddingGemma 2 handles retrieval and enrichment, according to Google.
What this means
The notable part of this release is the combination of a permissive license and a small footprint. At 740M parameters and a quantized multimodal footprint under 600MB of RAM (per Google), the model fits within the memory budget of current phones and laptops. A single embedding model covering audio and video also simplifies local retrieval pipelines, which currently often stack a speech-to-text model and a separate text embedder.
The Apache 2.0 license matters for builders. It permits commercial use and modification without the custom terms that have governed earlier Gemma releases, so developers can ship the model inside their own products.
The gaps limit any real assessment. Google has not published retrieval benchmark scores, context length or embedding dimensions in this announcement, so comparisons with existing open embedding models are not yet possible. The memory figures are also tied to one device and quantization setting. Teams evaluating the model should wait for the technical documentation and independent testing before committing to it.
Related Articles
Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0
Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says it runs in ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model on a quantized Pixel 11 Pro. Google also released a Mac app, AI Edge Foresight, to demonstrate it.
Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio
Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.
Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context
Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.
Mistral releases Large 4, a 1-trillion-parameter multimodal model, with open weights due in three weeks
Mistral AI released Mistral Large 4 (ML4), a multimodal model with one trillion parameters, on Tuesday. It is currently available only through a public guardrail endpoint, and Mistral plans to publish the weights in about three weeks after safety testing. Benchmark results, pricing and context window have not been disclosed.
Comments
Loading...