Google AI Edge Gallery launches on macOS with Gemma 4 12B, 12-billion-parameter model for local inference
Google launched AI Edge Gallery for macOS, allowing Mac users to run Google's Gemma models locally. The platform ships with five Gemma models, including the newly released Gemma 4 12B—a 12-billion-parameter multimodal model that handles text, vision, and audio while running on consumer laptops with 16GB of RAM.
Google AI Edge Gallery launches on macOS with Gemma 4 12B, 12-billion-parameter model for local inference
Google released AI Edge Gallery for macOS on June 3, 2026, enabling Mac users to run Google's Gemma models locally on their devices. The platform currently supports five models, with the flagship Gemma 4 12B representing Google's most capable consumer-focused local model to date.
Gemma 4 12B: specifications and capabilities
Gemma 4 12B is a 12-billion-parameter model that Google claims delivers performance comparable to its 26-billion-parameter mixture-of-experts model. According to Google, the model runs on consumer laptops with 16GB of RAM.
The model is multimodal, processing text, vision, and audio inputs. Google states it includes "good coding capabilities" for extracting insights from data on-device. No benchmark scores or pricing information have been disclosed, as the model is designed for local deployment rather than API access.
Available models
Google AI Edge Gallery for Mac currently offers access to five models, all instruction-tuned variants:
- Gemma-4-12B-it
- Gemma-4-E2B-it
- Gemma-4-E4B-it
- Gemma-3n-E2B-it
- Gemma-3n-E4B-it
Unlike platforms such as Ollama and LM Studio, which support thousands of open models from various providers, Google AI Edge Gallery is limited to Google's Gemma family.
Google AI Edge Eloquent
Google also launched AI Edge Eloquent for macOS, a dictation app that transcribes speech while editing for clarity and removing disfluencies. The app processes audio on-device and supports custom vocabulary for names and technical terms. Eloquent previously launched on iOS earlier in 2026.
What this means
Google is positioning itself as a competitor to established local inference platforms while maintaining a walled garden approach—only Google models are supported. The 12-billion-parameter size of Gemma 4 12B is notably larger than most consumer-focused local models, which typically range from 2 billion to 9 billion parameters. However, without published benchmarks or independent testing, claims about performance parity with larger models remain unverified. The 16GB RAM requirement makes Gemma 4 12B accessible to recent MacBook Air and MacBook Pro users, but excludes older hardware. Google's simultaneous release of a practical application (Eloquent) alongside developer tools (AI Edge Gallery) suggests a dual strategy targeting both general consumers and technical users.
Related Articles
OpenRouter Adds Auto-Updating Alias for Zhipu AI's GLM Flash Model Family
Z.ai has published GLM Flash Latest on OpenRouter, a routing alias that automatically points to the newest checkpoint in the GLM Flash lineup. It supports a 1.31M token context window and multimodal text, image, and video input at $0.07 per 1M input tokens and $0.25 per 1M output tokens.
Google Launches Pics, a Prompt-Based Design Tool Built Into Workspace
Google is launching Pics, an AI image creation and editing tool powered by its Nano Banana model, positioning it as a prompt-first alternative to Canva and Adobe Express. The tool rolls out first in Google Docs and Slides for Workspace and Google AI Pro/Ultra subscribers, with Drive support to follow.
Google Rolls Out Pics Image Editor to AI Pro, Ultra, and Business Workspace Subscribers
Google has begun broadly rolling out Pics, its AI image generation and editing app built on Nano Banana, to Google AI Pro and Ultra subscribers plus Business and Enterprise Workspace plans. The app supports object segmentation, text translation in 29 languages, and 4K upscaling, with direct integration into Docs and Slides.
Perplexity Launches Hybrid Compute for Mac, Splitting AI Tasks Between Cloud and Local Models
Perplexity's Mac app now supports Hybrid Compute, which starts tasks in the cloud and shifts sensitive steps to a local model running on-device. The feature requires Apple silicon with at least 24GB of unified memory and uses an open-sourced on-device PII classifier to mask private data before any cloud processing.
Comments
Loading...