Google releases Gemma 4 E2B, optimized to run natively on Pixel 10's Tensor G5 TPU
Google has released Gemma 4 E2B for TPU, a variant of its open-source Gemma 4 model optimized to run natively on the Tensor G5 chip in Pixel 10 devices. The multimodal model enables completely offline AI chat, image recognition, and audio transcription on Pixel 10, 10 Pro, 10 Pro XL, and 10 Pro Fold.
Google releases Gemma 4 E2B, optimized to run natively on Pixel 10's Tensor G5 TPU
Google announced Gemma 4 E2B for TPU today, a variant of its open-source Gemma 4 model designed to run natively on the Tensor Processing Unit in Pixel 10 devices. The announcement came at I/O Connect India, following a similar satellite event in Berlin last week.
Model specifications
Gemma 4 E2B runs on the Tensor G5's TPU and is supported on four devices: Pixel 10, 10 Pro, 10 Pro XL, and 10 Pro Fold. Google describes it as "state-of-the-art, powerful, yet remarkably lightweight," though specific parameter counts and benchmark scores were not disclosed.
The model is based on Gemma 4, which Google first introduced in April as the foundation for the upcoming Gemini Nano 4. Gemma is Google's series of open models designed for on-device execution.
Multimodal capabilities
Gemma 4 E2B supports three fully offline modes:
- AI Chat: On-device conversations with no internet connection required
- Ask Image: Object, plant, and issue identification from photos
- Ask Audio: Private audio transcription for lectures and notes
Google demonstrated "Mobile Actions" that allow users to control core phone functions like WiFi and maps through voice or text commands, all processed locally.
Real-world applications
Google highlighted two specific use cases:
Retail: Converting recipe ideas into localized in-store shopping maps completely offline, allowing customers to navigate stores without internet connectivity.
Automotive: Providing mechanics with immediate visual diagnostics from photos of faulty parts, enabling on-the-spot troubleshooting.
What this means
This release represents Google's push to run increasingly capable AI models entirely on-device, addressing privacy concerns and enabling functionality in areas with poor connectivity. By optimizing specifically for the Tensor G5's TPU architecture, Google is differentiating its Pixel hardware through exclusive AI capabilities that competitors cannot easily replicate. The retail and automotive examples suggest Google is targeting enterprise deployments where offline operation is critical, not just consumer use cases. However, without published benchmarks or comparisons to cloud-based alternatives, the actual performance trade-offs of this on-device approach remain unclear.
Related Articles
GLM-5.3-Flash Debuts as Zhipu AI's First Natively Multimodal Model, 320B Parameters with 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 while approaching Claude Opus 4.8 on coding and agentic benchmarks at a fraction of the cost. Unsloth has published quantized GGUF versions for local inference.
Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2
Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech
Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.
Comments
Loading...