model releaseGoogle DeepMind

NVIDIA releases Gemma 4 31B quantized model with 256K context, multimodal capabilities

TL;DR

NVIDIA has released a quantized version of Google DeepMind's Gemma 4 31B IT model, compressed to NVFP4 format for efficient inference on consumer GPUs. The 30.7B-parameter multimodal model supports 256K token context windows, handles text and image inputs with video frame processing, and maintains near-baseline performance across reasoning and coding benchmarks.

2 min read
0

NVIDIA Quantizes Google DeepMind's Gemma 4 31B for Efficient Inference

NVIDIA has released an NVFP4-quantized version of Google DeepMind's Gemma 4 31B IT model on Hugging Face, designed to run inference on consumer-grade NVIDIA GPUs while maintaining frontier-level performance for reasoning, coding, and multimodal tasks.

Model Specifications

The base Gemma 4 31B IT contains 30.7B parameters and supports a 256K-token context window—enabling extended document processing and multi-turn conversations. The model is multimodal, accepting text, image, and video inputs (up to 60 seconds at 1 FPS). It supports configurable visual token budgets (70, 140, 280, 560, 1120 tokens) and variable image aspect ratios. Vocabulary size is 262,144 tokens.

The model covers over 140 languages and uses a hybrid attention mechanism combining local sliding-window and global attention with Proportional RoPE for long-context stability.

Quantization Impact

NVIDIA's NVFP4 quantization (performed with nvidia-modelopt v0.42.0) shows minimal performance degradation:

  • GPQA Diamond: 75.71% → 75.46% (−0.25 points)
  • AIME 2025: 66.25% → 65.94% (−0.31 points)
  • MMLU Pro: 85.25% → 84.94% (−0.31 points)
  • LiveCodeBench (pass@1): 70.90% → 70.63% (−0.27 points)
  • Scicode (pass@1): 33.61% → 33.18% (−0.43 points)
  • Terminal-Bench Hard: 27.08% → 27.08% (no change)

The quantized model reduces memory requirements and enables deployment on NVIDIA Hopper architecture (H100) and newer Blackwell systems via vLLM.

Training Data and Licensing

The underlying Gemma 4 model was trained on large-scale multimodal data (text, code, images, audio) with a knowledge cutoff of January 2025. Google DeepMind applied CSAM filtering and safety processing. NVIDIA calibrated the quantized version using the CNN DailyMail dataset (300K+ articles).

The model is available under Apache License 2.0 with NVIDIA's Open Model License Agreement governing usage. It is cleared for both commercial and non-commercial use, though NVIDIA notes the model may amplify biases and toxicity from its training data.

Known Limitations

Gemma 4 31B IT can generate inaccurate information, omit key details, and produce toxic responses—particularly when prompted with adversarial inputs. The model does not blur or maintain aspect ratios of people, personal health information, or copyrighted content in images.

NVIDIA recommends developers validate the model against internal requirements for specific industries and use cases before deployment.

Deployment and Integration

The quantized model is optimized for vLLM inference engine with recommended tensor parallelism of 8 on H100 hardware. NVIDIA reports 6,292 downloads on Hugging Face in the past month, though no commercial pricing has been announced for hosted inference.

What This Means

This release makes frontier-class multimodal reasoning accessible on consumer and datacenter GPUs without API dependencies. The minimal performance loss (<0.5% on most benchmarks) validates NVFP4 quantization for production deployments. However, the lack of hosted inference offerings means organizations must self-host—requiring GPU infrastructure and operational overhead. The 256K context and multimodal capabilities position Gemma 4 31B as a direct competitor to larger proprietary models for organizations prioritizing deployment flexibility and cost control.

Related Articles

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev

Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.

Comments

Loading...