model releaseGoogle DeepMind

Google DeepMind Launches Gemini 3.8 Live, Claims #1 Spot on Speech-to-Speech Benchmark

TL;DR

Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-dialogue models that reason and execute background tasks without interrupting conversation. Google claims the Extended Thinking model ranks #1 on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6.

3 min read
0

Google DeepMind Ships Two New Voice Models

Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, positioning them as its most advanced real-time voice dialogue models to date. Both are available immediately through the Gemini API, Google AI Studio, the Gemini app, Search Live, and in private preview for Gemini Enterprise.

The two models target different use cases. Gemini 3.8 Live is built for low-latency, cost-efficient conversation with visual grounding, while Gemini 3.8 Live Extended Thinking is designed for multi-step reasoning tasks that require the model to "think" while continuing to speak.

Benchmark Claims

According to Google, Gemini 3.8 Live Extended Thinking took the top overall position on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. The company also claims the model leads in agentic task completion, scoring 68.6% on the τ-Voice benchmark and 35.1% on Sierra's τ-Voice-banking benchmark. On Big Bench Audio, Google reports a 97.7% score for reasoning capability.

Google states that the standard Gemini 3.8 Live model placed second in the Speech Agent Arena, a user-preference leaderboard, while remaining cheaper to run than the Extended Thinking variant. Google did not disclose specific pricing figures, parameter counts, or context window sizes for either model.

What's New

Gemini 3.8 Live processes visual input in near real time, letting the model incorporate what it sees into spoken responses. It automatically detects and switches between 97 supported languages mid-conversation without requiring a manual toggle. Both models can execute tool calls and API requests in the background — acknowledging a request verbally ("Let me check that…") and continuing dialogue while the task completes, rather than pausing the conversation.

Gemini 3.8 Live Extended Thinking adds live progress narration for multi-step tasks, verbally walking users through background work such as coordinating bookings, generating React components from sketches, or drafting business plans.

All audio output from the models is watermarked with SynthID, Google's imperceptible watermarking system intended to make AI-generated audio detectable.

Availability and Integration

Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, Search Live, and private preview for Gemini Enterprise. Gemini 3.8 Live Extended Thinking is available via the same channels plus Gemini Live for Google AI Pro and Ultra subscribers, and is rolling out to Google Workspace products including Docs, Gmail, and Keep. Developer platforms including Agora, LiveKit, Pipecat, Vercel, Fishjam, and Vision Agents have integrated the Gemini Live API to support real-time voice interfaces. Google also named Salesforce, Genspark, and Lumeris as enterprise partners evaluating the new models.

What This Means

The release pushes Google further into agentic voice interfaces — models that don't just transcribe and respond, but manage background tasks and tool calls while keeping a conversation active. The claimed benchmark leads on Artificial Analysis and τ-Voice are third-party evaluations, which gives them more weight than Google's own internal metrics, but the company has not published pricing or technical specs (context window, latency figures, parameter count), making it hard for developers to fully assess cost-effectiveness against competitors like OpenAI's Realtime API or Amazon's voice models. The real test will be whether the background task execution and multilingual switching hold up in production enterprise deployments rather than curated demos.

Related Articles

model release

Google Launches Gemini 3.8 Live, Undercutting OpenAI's GPT-Live-1 on Price by Up to 70%

Google DeepMind released Gemini 3.8 Live and a reasoning-enhanced Extended Thinking variant for voice agents, pricing audio input at $0.005/minute versus OpenAI's $0.05/minute for GPT-Live-1. The Extended Thinking model tops the Artificial Analysis Speech-to-Speech Leaderboard with 82.6 percent.

model release

Unverified 'GPT Astra' Model Appears on OpenRouter With 1.05M Token Context, No OpenAI Confirmation

OpenRouter is listing a model called 'OpenAI GPT Astra Latest' with a 1.05 million token context window and $10/$50 per-million-token pricing. OpenAI has made no public announcement, and the listing's own description says it is an auto-redirecting alias rather than a fixed model.

model release

OpenRouter Lists 'GPT Sol Latest' — An Alias Pointer to OpenAI's Newest Sol-Family Model, Not a Standalone Release

OpenRouter has added a listing called '~openai/gpt-sol-latest,' described as an alias that always points to the newest model in an undisclosed 'GPT Sol' family from OpenAI. The listing shows a 1050K token context window and pricing of $2.00 per million input tokens and $10.00 per million output tokens, but OpenAI has not publicly confirmed a model line by this name.

model release

OpenAI Launches GPT-Live-1 API for Full-Duplex Voice Apps That Talk and Listen Simultaneously

OpenAI has released GPT-Live-1 as a developer API, a speech model capable of full-duplex conversation—listening and talking simultaneously. It already powers ChatGPT's voice mode and costs $0.05 per minute, with benchmark scores showing sharp improvements over GPT-Realtime-2.1.

Comments

Loading...