product updateOpenAI

OpenAI Brings Agent-Controlling ChatGPT Voice Mode to Desktop App

TL;DR

OpenAI has brought ChatGPT Voice to its desktop app, letting users direct AI agents in ChatGPT Work and Codex through spoken commands. The feature, powered by OpenAI's new GPT-Live model family, can execute multi-step tasks like creating pull requests and debugging code.

3 min read
0

OpenAI has updated its ChatGPT desktop app to support ChatGPT Voice, giving users the ability to control AI agents and execute multi-step computer tasks using only spoken commands. The company announced the rollout Thursday, calling it a global release.

The desktop version is powered by GPT-Live, OpenAI's family of voice models launched earlier this month. According to OpenAI, ChatGPT Voice on desktop works with both ChatGPT Work and Codex, and can tap computer-use skills to search websites and navigate apps on a user's behalf.

On macOS, a companion feature called Appshots lets the app access what's currently on a user's screen, including alt-text, to inform its actions.

More capable than the mobile version

When ChatGPT Voice first launched on smartphones, it focused on natural conversation flow — smoother exchanges and better handling of interruptions — but according to OpenAI, it could not take actions on the phone itself. Thursday's desktop update expands on that foundation, allowing users to dictate complex, multi-step commands and respond in real time when the agent needs clarification or input.

In a demo video posted to X, OpenAI showed a developer issuing a single voice command that instructed ChatGPT to create a new thread, open a pull request, and identify the root cause of a bug — three distinct actions triggered from one spoken instruction.

OpenAI also said users can access ChatGPT Voice within Codex from the iOS app through remote access, extending the desktop-class agent control to mobile in a limited form.

Pricing for ChatGPT Voice was not disclosed in OpenAI's announcement, and the company did not specify whether the feature requires a ChatGPT Plus, Team, or Enterprise subscription beyond existing ChatGPT Work and Codex access.

Competitive pressure from Anthropic

The update lands the same week Anthropic refreshed its own voice mode for Claude, which now taps the company's Opus, Sonnet, and Haiku models to complete tasks inside third-party apps including Gmail, Calendar, Slack, Notion, and Canva. Anthropic separately launched Opus 5 within the same news cycle, according to TechCrunch's coverage.

Both companies are racing to position voice as the primary interface for agentic computer control rather than a novelty for casual conversation. OpenAI's framing — voice as a control layer for coding agents and workplace tools — signals it sees developers and enterprise users, not just consumers, as the target audience for GPT-Live.

What this means

Voice interfaces are shifting from conversational novelty to functional control surfaces for AI agents. OpenAI's move to let developers dictate multi-step coding tasks — create a thread, open a PR, find a bug — by voice suggests the company is betting that hands-free agent orchestration will matter more for retention than chat quality alone. The near-simultaneous Anthropic voice update for Claude, tied to workplace apps like Gmail and Slack, confirms this is now a two-horse race over who owns the voice-to-agent layer in daily computer use. Neither company has disclosed pricing specifics for these voice features, making it hard to assess whether this becomes a paid differentiator or a bundled perk for existing subscribers. Expect scrutiny over reliability next: multi-step voice commands that touch code repositories or business apps raise the stakes for error correction and undo mechanisms in ways that text-based chat never did.

Related Articles

product update

Google's Gemini Spark Agent Gains Ability to Manage Google Photos Libraries

Google announced that Gemini Spark, its personal AI agent, can now execute tasks directly inside Google Photos—editing images, curating albums, and turning flyers into calendar events. The rollout begins over the next few weeks for Gemini AI Pro and Ultra subscribers in the U.S.

benchmark

Simon Willison's Pelican Benchmark Shows GPT-6 Astra Outperforming GPT-5.6 Sol at Every Reasoning Level

Developer Simon Willison ran his signature 'pelican riding a bicycle' SVG test on newly-accessed GPT-6 Astra across five reasoning levels, comparing results against GPT-5.6 Sol, Terra, and Luna. Even Astra's lowest reasoning setting reportedly beat every Sol output, though Astra costs roughly twice as much per token.

product update

AWS Publishes Reference Architecture for Multimodal WhatsApp Ordering Agents Using Bedrock AgentCore and Nova 2

AWS published a reference architecture showing how to deploy a WhatsApp ordering assistant on Amazon Bedrock AgentCore, using Nova 2 Lite for text and Nova 2 Sonic for voice, with shared cross-channel memory and MCP-based tool access to backend systems.

product update

Gemini Overlay on Android Adds Minimize Button for Multitasking Bubble

Google is widely rolling out a new Minimize button for the Gemini overlay on Android, which collapses conversations into a floating bubble users can drag or tap to expand. The feature currently supports six fixed positions and has a reported bug that resets bubble placement after each minimization.

Comments

Loading...