product updateOpenAI

OpenAI Brings Agent-Controlling ChatGPT Voice Mode to Desktop App

TL;DR

OpenAI has brought ChatGPT Voice to its desktop app, letting users direct AI agents in ChatGPT Work and Codex through spoken commands. The feature, powered by OpenAI's new GPT-Live model family, can execute multi-step tasks like creating pull requests and debugging code.

3 min read
0

OpenAI has updated its ChatGPT desktop app to support ChatGPT Voice, giving users the ability to control AI agents and execute multi-step computer tasks using only spoken commands. The company announced the rollout Thursday, calling it a global release.

The desktop version is powered by GPT-Live, OpenAI's family of voice models launched earlier this month. According to OpenAI, ChatGPT Voice on desktop works with both ChatGPT Work and Codex, and can tap computer-use skills to search websites and navigate apps on a user's behalf.

On macOS, a companion feature called Appshots lets the app access what's currently on a user's screen, including alt-text, to inform its actions.

More capable than the mobile version

When ChatGPT Voice first launched on smartphones, it focused on natural conversation flow — smoother exchanges and better handling of interruptions — but according to OpenAI, it could not take actions on the phone itself. Thursday's desktop update expands on that foundation, allowing users to dictate complex, multi-step commands and respond in real time when the agent needs clarification or input.

In a demo video posted to X, OpenAI showed a developer issuing a single voice command that instructed ChatGPT to create a new thread, open a pull request, and identify the root cause of a bug — three distinct actions triggered from one spoken instruction.

OpenAI also said users can access ChatGPT Voice within Codex from the iOS app through remote access, extending the desktop-class agent control to mobile in a limited form.

Pricing for ChatGPT Voice was not disclosed in OpenAI's announcement, and the company did not specify whether the feature requires a ChatGPT Plus, Team, or Enterprise subscription beyond existing ChatGPT Work and Codex access.

Competitive pressure from Anthropic

The update lands the same week Anthropic refreshed its own voice mode for Claude, which now taps the company's Opus, Sonnet, and Haiku models to complete tasks inside third-party apps including Gmail, Calendar, Slack, Notion, and Canva. Anthropic separately launched Opus 5 within the same news cycle, according to TechCrunch's coverage.

Both companies are racing to position voice as the primary interface for agentic computer control rather than a novelty for casual conversation. OpenAI's framing — voice as a control layer for coding agents and workplace tools — signals it sees developers and enterprise users, not just consumers, as the target audience for GPT-Live.

What this means

Voice interfaces are shifting from conversational novelty to functional control surfaces for AI agents. OpenAI's move to let developers dictate multi-step coding tasks — create a thread, open a PR, find a bug — by voice suggests the company is betting that hands-free agent orchestration will matter more for retention than chat quality alone. The near-simultaneous Anthropic voice update for Claude, tied to workplace apps like Gmail and Slack, confirms this is now a two-horse race over who owns the voice-to-agent layer in daily computer use. Neither company has disclosed pricing specifics for these voice features, making it hard to assess whether this becomes a paid differentiator or a bundled perk for existing subscribers. Expect scrutiny over reliability next: multi-step voice commands that touch code repositories or business apps raise the stakes for error correction and undo mechanisms in ways that text-based chat never did.

Related Articles

changelog

OpenAI Publishes GPT-6 Astra Prompting Guide With Banned 'Slop Words' List

OpenAI has published detailed prompting guidance for GPT-6 Astra, addressing the model's tendency to over-clarify, over-test, and use clichéd AI phrasing. The documentation includes specific prompts to encourage more autonomous action and a blocklist of banned words and phrases.

model release

OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol

OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.

product update

Google's Gemini Spark Agent Gains Ability to Manage Google Photos Libraries

Google announced that Gemini Spark, its personal AI agent, can now execute tasks directly inside Google Photos—editing images, curating albums, and turning flyers into calendar events. The rollout begins over the next few weeks for Gemini AI Pro and Ultra subscribers in the U.S.

benchmark

Unverified 'GPT-6 Astra' Reportedly Completes Portal Solo in Under 24 Hours, No Official OpenAI Confirmation

A developer named cozyblaze posted on X that a model called 'GPT-6 Astra' completed Portal from start to finish without human intervention in 23 hours 43 minutes. OpenAI has not confirmed the existence of a model by that name, and all details come from a single third-party account.

Comments

Loading...