product update

Gemini for macOS Rolls Out Voice Control With Screen-Aware Task Execution

TL;DR

Google is rolling out advanced voice control to Gemini for macOS version 1.88, combining Gboard Rambler-style dictation with a screen-aware assistant that can summarize files, rewrite text, and generate images by voice. The feature, previewed at I/O 2026, activates via a long-press of the Fn key and requires Gemini reasoning to be enabled for the advanced capabilities.

2 min read
0

Google has begun rolling out advanced voice control to Gemini for macOS, bringing dictation quality comparable to Gboard Rambler alongside a new screen-aware assistant mode. The feature was previewed at I/O 2026 in May and is now shipping in version 1.88 of the Gemini macOS app.

What's new

The update centers on two capabilities Google describes as letting users "speak naturally into any window on your desktop."

The first is intelligent dictation. According to Google, Gemini transcribes spoken words into "clean, polished text," automatically stripping filler words like "umm" and "ah" and accounting for mid-sentence corrections to capture intended meaning. The formatted output is inserted directly at the current cursor position, mirroring the transcription approach used by Gboard Rambler on upcoming Gemini Intelligence phones.

The second capability is screen-context understanding, which lets Gemini interpret what's currently displayed on a user's desktop to execute more complex requests. This mode is opt-in and requires enabling Gemini reasoning in the app's settings. Google outlines three use cases:

  • Extract and summarize information: Users can highlight local files, images, or documents and issue voice commands such as "Read these vet files and summarize my dog's medical history in an email to the kennel."
  • Compose and rewrite text: Highlighting text anywhere on screen allows voice-driven rewrites, tone adjustments, and reformatting — for example, "Turn these notes into an executive summary with a TL;DR at the top."
  • Generate and edit images: Users can create or iterate on visuals by voice, referencing an existing image on screen, such as requesting a "dark-mode version" of an illustration.

How to activate it

Users can trigger voice control by long-pressing the Fn key anywhere in macOS, or by tapping a new screen-sharing button at the end of the "Ask Gemini" prompt box. Once active, a floating pill containing a waveform appears at the bottom of the screen to indicate Gemini is listening.

The rollout is global but currently limited to English, with Google stating additional languages are "coming soon." Users need version 1.88 or later of Gemini for macOS to access the feature.

What this means

This update pushes Gemini further into direct competition with OS-level assistants like Apple's Siri and Microsoft's Copilot, both of which have struggled to ship reliable screen-context and voice-command features on desktop. By tying dictation quality to the same engine powering Gboard Rambler on Android, Google is signaling a unified voice architecture across its consumer AI products rather than platform-specific implementations.

The gated nature of the screen-aware mode — requiring Gemini reasoning to be manually enabled — suggests Google is still cautious about the reliability or compute cost of letting a model parse arbitrary desktop content on command. The English-only limitation at launch also indicates this is an early-stage rollout rather than a finished product, with broader language support pending before Google likely positions this as a mainstream macOS feature rather than a beta capability confined to Gemini app users.

Related Articles

product update

Gemini for macOS Rolls Out Voice Control With Screen-Aware Dictation and Editing

Google is rolling out advanced voice control for Gemini on macOS, version 1.88, letting users dictate cleaned-up text and issue voice commands that reference on-screen content. The feature, previewed at I/O 2026, requires opt-in Gemini reasoning to unlock summarization, rewriting, and image editing tasks.

product update

Google Simplifies Gemini App's Thinking Level Picker, Adds Notification Controls

Google is simplifying the Gemini app's model picker by collapsing the two-stage 'Standard' vs 'Extended thinking' selector into a single toggle. The company is also rolling out new notification settings on Android and reorganizing the Gemini Spark task interface.

product update

Google Expands Gemini Spark Agentic Assistant to All AI Pro and Ultra Subscribers

Google is expanding access to Gemini Spark, its agentic AI assistant built on Gemini 3.5, to all Google AI Pro subscribers in the US and Google AI Ultra subscribers globally. The rollout excludes free-tier users and, for Ultra, customers in the EEA, Switzerland, the UK, and Nigeria.

product update

Replit Launches 'Replit Design,' an AI Design Suite Powered by Claude, GPT-5, Gemini, Kimi, and GLM

Replit has launched Replit Design, a browser-based AI design suite that lets users generate apps, sites, and brand assets using models including Claude, GPT-5, Gemini, Kimi, and GLM. The product replaces Replit's earlier Canvas tool and integrates the Mobbin UI reference library directly into the workflow.

Comments

Loading...