Google DeepMind releases Gemma 4, open multimodal models with 256K context and reasoning
Google DeepMind has released Gemma 4, a family of open-weights multimodal models ranging from 2.3B to 31B parameters with support for text, images, video, and audio. The models feature context windows up to 256K tokens, built-in reasoning modes, and native function calling for agentic workflows.
Gemma 4 31B Instruct — Quick Specs
Google DeepMind Releases Gemma 4: Open Multimodal Models with Extended Context
Google DeepMind has released Gemma 4, a family of open-weights models spanning from 2.3B to 31B parameters with multimodal capabilities and extended context windows up to 256K tokens. The release includes both dense and Mixture-of-Experts (MoE) architectures designed for deployment across devices from mobile phones to data center servers.
Model Sizes and Specifications
Gemma 4 offers four distinct variants:
- E2B: 2.3B effective parameters (5.1B with embeddings), 128K context, text/image/audio support
- E4B: 4.5B effective parameters (8B with embeddings), 128K context, text/image/audio support
- 26B A4B (MoE): 25.2B total parameters with 3.8B active parameters, 256K context, text/image support
- 31B Dense: 30.7B parameters, 256K context, text/image support
The smaller E2B and E4B models use Per-Layer Embeddings (PLE) technology to reduce effective parameter counts, enabling efficient deployment on edge devices. The 26B A4B variant uses a Mixture-of-Experts approach with 128 total experts and 8 active experts, claiming inference speeds comparable to a 4B model while maintaining 26B total capacity.
Capabilities and Architecture
All Gemma 4 models support text and image inputs with variable aspect ratios and resolutions. The E2B and E4B models additionally include native audio support with automatic speech recognition and multilingual speech-to-translation capabilities. Video understanding is available through frame sequence processing.
Key features include:
- Reasoning: Configurable thinking modes enabling step-by-step reasoning before response generation
- Function Calling: Native support for structured tool use and agentic workflows
- Hybrid Attention: Combines local sliding window attention with full global attention, with Proportional RoPE optimization for memory efficiency
- Multilingual: Pre-trained on 140+ languages with out-of-the-box support for 35+
- Native System Prompt Support: Structured conversation control
Benchmark Performance
The instruction-tuned models show significant improvements in reasoning and coding tasks:
Gemma 4 31B achieves:
- MMLU Pro: 85.2%
- AIME 2026 (no tools): 89.2%
- LiveCodeBench v6: 80.0%
- Codeforces ELO: 2150
- GPQA Diamond: 84.3%
- MATH-Vision: 85.6%
- Long Context MRCR v2 (128K needle): 66.4%
Gemma 4 26B A4B demonstrates strong performance-to-efficiency trade-offs:
- MMLU Pro: 82.6%
- AIME 2026 (no tools): 88.3%
- LiveCodeBench v6: 77.1%
- Codeforces ELO: 1718
Smaller models show corresponding improvements over Gemma 3 27B, with E2B scoring 60.0% on MMLU Pro compared to Gemma 3's 67.6% baseline.
Release Details
The models are released under Apache 2.0 licensing as both pre-trained and instruction-tuned variants. Unsloth has released GGUF quantized versions optimized for local inference. The models are available through Hugging Face with support for the latest Transformers library.
Google DeepMind emphasizes on-device deployment viability for the smaller models while positioning larger variants for consumer GPU and server deployment. The hybrid architecture and context window scaling address trade-offs between inference speed and reasoning depth for long-context tasks.
What this means
Gemma 4 represents a significant shift toward production-ready open models with genuine multimodal capabilities and reasoning support at multiple scale points. The MoE variant offers a novel efficiency approach for teams balancing model capacity with inference latency constraints. Notably absent from the release are specific pricing details for cloud inference—unlike proprietary alternatives—since these are open-weights models suitable for self-hosted deployment. The 256K context window and strong long-context benchmark performance position these models competitively for document analysis and extended reasoning tasks against closed commercial alternatives.
Related Articles
Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta
Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.
inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter
inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter
Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.
Comments
Loading...