Gemini 3.1 Flash Live scores 95.9% on Big Bench Audio, Google's fastest voice model
Google has released Gemini 3.1 Flash Live, its new voice and audio AI model, scoring 95.9% on the Big Bench Audio Benchmark at high thinking levels—second only to Step-Audio R1.1 Realtime at 97.0%. Response times range from 0.96 seconds at minimal thinking to 2.98 seconds at high thinking, with pricing held at $0.35 per hour of audio input and $1.40 per hour of audio output.
Gemini 3.1 Flash Live — Quick Specs
Google Releases Gemini 3.1 Flash Live Voice Model
Google has unveiled Gemini 3.1 Flash Live, a new voice and audio AI model positioned as the company's best-performing audio offering to date. The model is now available through the Gemini Live API, Google AI Studio, Gemini Live, and Search Live across over 200 countries.
Performance and Capabilities
According to Artificial Analysis benchmarking, Gemini 3.1 Flash Live achieves 95.9% on the Big Bench Audio Benchmark when configured to its "High" thinking level, placing it second only to Step-Audio R1.1 Realtime, which scores 97.0%. At the "Minimal" thinking level, the model's score drops to 70.5% but response time improves significantly.
Response latency varies by configuration:
- High thinking: 2.98-second response time
- Minimal thinking: 0.96-second response time
Google claims the model delivers improved pitch and emotion detection compared to its predecessor, with enhanced reliability in noisy environments. Developers can now configure thinking levels directly, allowing trade-offs between output quality and latency.
Pricing
Gemini 3.1 Flash Live maintains identical pricing to its Gemini 2.5 predecessor:
- Audio input: $0.35 per hour
- Audio output: $1.40 per hour
Google positions this as among the cheapest audio AI models available. For comparison, Step-Audio R1.1 Realtime offers lower input pricing but charges more for audio output.
Deployment and Integration
The model now powers live mode functionality within the Gemini app, enabling real-time voice conversations. Integration is available through multiple access points, supporting developers building voice applications across the Google ecosystem.
What this means
Gemini 3.1 Flash Live competes directly with Step-Audio R1.1 Realtime in the high-performance voice AI space, with nearly matching benchmark scores at a lower price point. The configurable thinking levels provide developers genuine flexibility for latency-sensitive applications—a meaningful improvement over fixed-performance models. At 0.96 seconds for minimal thinking, the model targets real-time conversational use cases where sub-second response times matter. The widespread availability across 200+ countries and multiple access methods signals Google's commitment to voice as a core interaction paradigm for Gemini products.
Related Articles
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks
Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Black Forest Labs Launches FLUX 3 Video, Claims It Beats Seedance 2.0 on Elo Rankings
Black Forest Labs has made FLUX 3 Video generally available via its API, offering up to 20-second HD/Full HD clips with native audio and lip-sync in 14+ languages. The company claims its internal Elo benchmarks put the model ahead of Seedance 2.0, Gemini Omni Flash, and Minimax H3.
Comments
Loading...