voice AI

19 articles tagged with voice AI

September 20, 2026
researchTencent

Tencent Unveils Gander, a Voice AI That Keeps Talking While a Separate 'Brain' Handles Background Tasks

Tencent's Hunyuan Speech team, working with university researchers, has released a technical report on Gander, a voice AI model that separates real-time conversation handling from complex background reasoning. The model interrupts users less often than GPT-Realtime, Gemini Live, and Grok in tests, but lags on task accuracy and video/audio understanding.

September 17, 2026
product update

Instinct and Meta's Muse AI Agents Both Add Phone-Calling Features Within Hours of Each Other

Instinct and Meta's Muse AI assistants both rolled out phone-calling capabilities on September 16, 2026, letting agents book restaurants, manage service calls, and act as a concierge. Instinct is reportedly in talks to raise $1 billion at a $10 billion valuation.

September 15, 2026
model release

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models, Tops Speech-to-Speech Benchmark

Google has announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new voice dialogue models that claim the #1 spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. The models are rolling out to Gemini Live and power new conversational features in Gmail, Docs, and Keep.

model releaseGoogle DeepMind

Google Launches Gemini 3.8 Live, Undercutting OpenAI's GPT-Live-1 on Price by Up to 70%

Google DeepMind released Gemini 3.8 Live and a reasoning-enhanced Extended Thinking variant for voice agents, pricing audio input at $0.005/minute versus OpenAI's $0.05/minute for GPT-Live-1. The Extended Thinking model tops the Artificial Analysis Speech-to-Speech Leaderboard with 82.6 percent.

model releaseGoogle DeepMind

Google DeepMind Launches Gemini 3.8 Live, Claims #1 Spot on Speech-to-Speech Benchmark

Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-dialogue models that reason and execute background tasks without interrupting conversation. Google claims the Extended Thinking model ranks #1 on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6.

September 10, 2026
model releaseOpenAI

OpenAI Launches GPT-Live-1 API for Full-Duplex Voice Apps That Talk and Listen Simultaneously

OpenAI has released GPT-Live-1 as a developer API, a speech model capable of full-duplex conversation—listening and talking simultaneously. It already powers ChatGPT's voice mode and costs $0.05 per minute, with benchmark scores showing sharp improvements over GPT-Realtime-2.1.

August 26, 2026
model release

Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing

Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.

model release

Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler

Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.

August 10, 2026
product updateNVIDIA

NVIDIA Releases Magpie TTS Multilingual Update: 364M-Parameter Open-Weights Model Now Supports 12 Languages, Sub-50ms La

NVIDIA's Magpie TTS Multilingual, a 364M-parameter open-weights text-to-speech model, now supports 12 languages after adding Modern Standard Arabic, Korean, and Brazilian Portuguese. The model achieves 32ms time-to-first-audio on B200 GPUs and improves speech quality across French, Spanish, and German.

August 5, 2026
model releaseNVIDIA

NVIDIA Releases Nemotron VoiceChat 11B, an Open Full-Duplex Speech Model with Live Tool Calling

NVIDIA has released NemotronLabs VoiceChat 11B, an 11-billion-parameter end-to-end full-duplex speech model that unifies streaming speech understanding and generation in one architecture. The model claims to be the first open full-duplex system to support live tool calling during natural conversation, with ~450ms turn-taking latency.

July 16, 2026
product updateAmazon Web Services

AWS launches AgentCore platform for building voice AI agents with Amazon Nova 2 Sonic

AWS has released AgentCore, a new platform for hosting and running voice-based AI agents, integrated with Amazon Nova 2 Sonic for real-time speech capabilities. The platform uses the open Model Context Protocol (MCP) to connect agents to backend systems and deploys each conversation in isolated microVMs.

July 9, 2026
product updateOpenAI

OpenAI launches GPT-Live voice models with full-duplex conversation and simultaneous web search

OpenAI has released new voice models for ChatGPT that use full-duplex architecture, allowing the AI to speak and listen simultaneously. GPT-Live-1 is available for paid subscribers (Go, Plus, Pro), while GPT-Live-1 mini serves free users.

June 24, 2026
product updateAmazon Web Services

Loka Achieves 87% Speech Reasoning Accuracy Using Amazon Nova 2 Sonic, Outperforming GPT Realtime and Gemini

Loka built a conversational voice agent using Amazon Nova 2 Sonic that achieved 87.0% speech reasoning accuracy on Big Bench Audio, surpassing GPT Realtime at 83.0% and Gemini 2.5 Flash Native Audio at 71.0%. The system delivers Time to First Audio of 1.39 seconds at approximately $0.27 per hour of input audio.

May 19, 2026
product update

Google launches Gmail Live, voice-powered AI inbox assistant for Ultra subscribers this summer

Google announced Gmail Live at IO 2026, a Gemini-powered conversational AI feature that allows users to ask natural language questions about their inbox instead of typing search terms. The voice-powered tool will roll out this summer exclusively to Google AI Ultra subscribers.

product update

Google adds voice prompting to Docs, Keep, and Gmail via Gemini AI

Google unveiled voice-based prompting for Docs, Keep, and Gmail at I/O 2026, powered by Gemini AI. The features enable document creation, note organization, and email search through spoken commands, launching this summer for Google AI Premium subscribers and Workspace business users.

May 12, 2026
model release

Mira Murati's Thinking Machines announces full-duplex AI model with 0.40-second response time

Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, announced TML-Interaction-Small, a full-duplex AI model that processes input while generating responses simultaneously. The company claims 0.40-second response time, matching natural human conversation speed.

May 7, 2026
product updateOpenAI+1

OpenAI launches GPT-Realtime-2 with GPT-5-class reasoning, adds real-time translation across 70 languages

OpenAI has added three voice intelligence features to its Realtime API: GPT-Realtime-2 with GPT-5-class reasoning for complex conversational requests, GPT-Realtime-Translate supporting 70 input languages and 13 output languages, and GPT-Realtime-Whisper for live speech-to-text transcription. Translation and transcription are billed by the minute, while GPT-Realtime-2 uses token-based pricing.

model releaseOpenAI

OpenAI releases GPT-Realtime-2 reasoning voice model with two specialized variants for translation and transcription

OpenAI has released three new realtime voice models through its Realtime API: GPT-Realtime-2 with GPT-5-class reasoning capabilities, GPT-Realtime-Translate supporting 70 input languages, and GPT-Realtime-Whisper for streaming transcription. The models are priced at $32-64 per 1M audio tokens for GPT-Realtime-2, and $0.017-0.034 per minute for the specialized variants.

April 30, 2026
product update

Google deploys Gemini AI to millions of existing cars, replacing Google Assistant

Google announced it will deploy Gemini AI to vehicles with Google built-in, replacing the current Google Assistant. General Motors confirmed 4 million vehicles from model year 2022 and newer across Cadillac, Chevrolet, Buick, and GMC brands will receive the update, with the rollout beginning in the U.S. with English-language support.