Gemma 4 12B

Google DeepMind🇺🇸 United States
active

Version History

4-12Bmajor

Gemma 4 12B is a new 12-billion-parameter multimodal model designed for local inference on consumer laptops. Google claims it matches the performance of their 26-billion-parameter mixture-of-experts model while running on devices with 16GB RAM.

4-12bmajor

First mid-sized Gemma model with native audio support, eliminates multimodal encoders for direct vision and audio processing through LLM backbone. Requires only 16GB RAM for local inference.

Benchmark Scores

Full leaderboard →
77.5%
AIME 2026
75.3%
GPQA
73.4%
Hallucination Rate
77.2%
MMLU-Pro
127.0 tokens_per_sec
Speed (tok/s)

Coverage