Google releases Gemini 3.1 Flash Live, its highest-quality audio model for real-time voice AI
Google has released Gemini 3.1 Flash Live, its highest-quality audio and voice model designed for real-time dialogue. The model scores 90.8% on ComplexFuncBench Audio and 36.1% on Scale AI's Audio MultiChallenge with reasoning enabled, with improved tonal understanding and lower latency compared to previous versions.
Gemini 3.1 Flash Live — Quick Specs
Google Releases Gemini 3.1 Flash Live, Its Highest-Quality Audio Model
Google has launched Gemini 3.1 Flash Live, a real-time audio and voice model designed to deliver more natural and reliable voice interactions. The model is now available to developers via the Gemini Live API in Google AI Studio, to enterprises through Gemini Enterprise for Customer Experience, and to all users via Gemini Live and Search Live.
Performance Benchmarks
On ComplexFuncBench Audio—which measures multi-step function calling with various constraints—Gemini 3.1 Flash Live achieves 90.8%, outperforming the previous model. On Scale AI's Audio MultiChallenge, which tests complex instruction following and real-world audio conditions including interruptions and hesitations, the model scores 36.1% with "thinking" mode enabled.
Google claims the model delivers improved latency compared to its predecessor, enabling faster response times for voice-first applications. The company also reports enhanced tonal understanding, allowing the model to recognize acoustic nuances like pitch and pace, and to dynamically adjust responses based on user expressions of frustration or confusion.
Developer Features
For developers, Gemini 3.1 Flash Live enables building voice agents capable of executing complex, multi-step tasks in noisy environments. The model supports function calling with improved reliability at scale. In Gemini Live, users can maintain conversation context for twice as long as with the previous model, preserving continuity during extended brainstorming sessions.
Companies including Verizon, LiveKit, and The Home Depot have provided positive feedback on the model's performance in production workflows, highlighting natural conversation quality.
Multilingual and Global Rollout
Gemini 3.1 Flash Live is inherently multilingual, enabling this week's global expansion of Search Live to over 200 countries and territories. Users can now conduct real-time, multimodal conversations with Google Search in their preferred language.
Safety and Watermarking
All audio generated by Gemini 3.1 Flash Live is watermarked using Google's SynthID technology. According to Google, this imperceptible watermark is embedded directly into audio output, enabling reliable detection of AI-generated content to help prevent misinformation.
What This Means
Gemini 3.1 Flash Live represents a meaningful advancement in real-time voice AI, with concrete benchmark improvements in function calling and instruction following. The model's expansion to 200+ countries positions Google to compete more aggressively in voice-first AI interfaces. The SynthID watermarking approach addresses growing regulatory and safety concerns around synthetic audio detection. For enterprises and developers, the improved tonal understanding and lower latency reduce friction in deploying voice agents for customer service and complex task automation.
Related Articles
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Anthropic Launches Claude Opus 5.5 at 20% Lower List Price, Claims Parity with Claude Fable 5.1
Anthropic released Claude Opus 5.5, the first model in its new 5.5 family, cutting list pricing 20% to $4/$20 per 1M input/output tokens while claiming performance on par with Claude Fable 5.1. Independent analysis shows the cost savings largely disappear at maximum reasoning effort due to higher token consumption.
Anthropic Ships Claude Opus 5.5, OpenAI Counters with GPT-6 Sol and Luna Hours Later, Triggering Sharp Price Cuts
Anthropic released Claude Opus 5.5 with a 20% price cut, and roughly an hour later OpenAI shipped GPT-6 Sol and GPT-6 Luna at roughly half the price of their GPT-5.6 predecessors. The releases follow Grok 4.7 and MiMo v2.6 from the previous day, intensifying competition among frontier model providers.
Comments
Loading...