Google Launches Gemini 3.8 Live and Extended Thinking Voice Models, Tops Speech-to-Speech Benchmark
Google has announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, new voice dialogue models that claim the #1 spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. The models are rolling out to Gemini Live and power new conversational features in Gmail, Docs, and Keep.
Google today announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, positioning them as the company's "most advanced live dialogue models yet." The models succeed Gemini 3.1 Flash Live, which launched in March 2026, and are rolling out today to Gemini Live along with new conversational features across Gmail, Docs, and Keep.
Benchmark claims
According to Google, Gemini 3.8 Live Extended Thinking holds the #1 overall position on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. The company also claims leading agentic task performance with 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark. On reasoning, Google reports a 97.7% score on Big Bench Audio.
The base Gemini 3.8 Live model, described as built for "scale and cost efficiency," claims second place in the Speech Agent Arena. Pricing for either model has not yet been disclosed.
What's new
Gemini 3.8 Live Extended Thinking is designed for high-complexity tasks and "reasons and speaks simultaneously," according to Google. The model uses verbal filler cues — such as "Let me check that…" — to acknowledge prompts while background processing occurs, and narrates progress during multi-step tasks. Google says this is intended to preserve conversational flow rather than forcing users to wait silently during longer operations.
The model powers three newly launched conversational products: Gmail Live (conversational search within email), Docs Live (draft generation and editing), and Keep Live (note creation). It is also beginning rollout to the standard Gemini Live app today.
The base Gemini 3.8 Live model focuses on "conversational intelligence with fluid dialogue and visual grounding," per Google. It processes visual inputs in near real-time and can execute tool and API calls in the background while conversation continues uninterrupted. The model supports mid-conversation language switching across 97 languages. Gemini 3.8 Live will also power the Search Live experience within Google's AI Mode.
Google has not disclosed parameter counts, training cutoff date, or context window size for either model.
What this means
This release signals Google's push to make voice AI feel less like a request-response system and more like a continuous conversation partner — a design pattern OpenAI and others have also pursued with real-time voice models. The emphasis on "speaking while thinking" and background task execution addresses a known weakness in voice assistants: dead air during processing breaks conversational immersion.
The benchmark claims, particularly the top Speech to Speech Quality Index ranking, are self-reported by Google and have not been independently verified by Artificial Analysis or third parties as of publication. The second-place finish in Speech Agent Arena for the base model, rather than first, suggests Google is positioning Extended Thinking as its flagship voice product while the standard model competes on cost and scale.
Deployment across Gmail, Docs, and Keep signals Google's strategy of embedding voice-driven AI directly into its productivity suite rather than treating it as a standalone assistant feature — a move that could pressure Microsoft's Copilot voice integrations if user adoption follows.
Related Articles
Google Launches Gemini 3.8 Live, Undercutting OpenAI's GPT-Live-1 on Price by Up to 70%
Google DeepMind released Gemini 3.8 Live and a reasoning-enhanced Extended Thinking variant for voice agents, pricing audio input at $0.005/minute versus OpenAI's $0.05/minute for GPT-Live-1. The Extended Thinking model tops the Artificial Analysis Speech-to-Speech Leaderboard with 82.6 percent.
Google DeepMind Launches Gemini 3.8 Live, Claims #1 Spot on Speech-to-Speech Benchmark
Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-dialogue models that reason and execute background tasks without interrupting conversation. Google claims the Extended Thinking model ranks #1 on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6.
OpenAI Launches GPT-Live-1 API for Full-Duplex Voice Apps That Talk and Listen Simultaneously
OpenAI has released GPT-Live-1 as a developer API, a speech model capable of full-duplex conversation—listening and talking simultaneously. It already powers ChatGPT's voice mode and costs $0.05 per minute, with benchmark scores showing sharp improvements over GPT-Realtime-2.1.
AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks
Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.
Comments
Loading...