Google releases Gemini 3.1 Flash Live, claims improved audio recognition and lower latency for voice conversations
Google announced Gemini 3.1 Flash Live as its updated audio and voice model for Gemini Live and Search Live. The model claims improved acoustic recognition, better background noise filtering, support for over 90 languages, and lower latency compared to 2.5 Flash Native Audio.
Gemini 3.1 Flash Live — Quick Specs
Google announced Gemini 3.1 Flash Live today as an upgrade to its audio and voice capabilities for Gemini Live and Search Live, now available in preview via the Gemini Live API in Google AI Studio.
According to Google, 3.1 Flash Live is the company's "highest-quality audio and voice model yet," with specific improvements in acoustic processing. The model claims to be "more effective at recognizing acoustic nuances like pitch and pace" and includes enhanced background noise filtering that better "discerns relevant speech from environmental sounds like traffic or television."
Key Technical Claims
Google claims the following improvements:
- Language support: Over 90 languages for real-time multi-modal conversations
- Latency: Lower latency compared to 2.5 Flash Native Audio
- Conversation length: On Android and iOS, can "follow the thread of your conversation for twice as long"
- Tool integration: "Significantly improved the model's ability to trigger external tools and deliver information during live conversations"
- Instruction adherence: Better compliance with complex system instructions, maintaining "operational guardrails even when conversations take unexpected turns"
- Response quality: Faster responses with "fewer awkward pauses" and dynamic adjustment of answer length and tone
Search Live Expansion
Google is deploying Gemini 3.1 Flash Live to roll out Search Live globally across over 200 countries and all languages where AI Mode is currently available. This includes audio and video (Google Lens) capabilities for back-and-forth conversations with Google Search.
The company claims that on Gemini Live, the new model delivers faster responses and can maintain conversation context for longer periods, which Google describes as "keeping your train of thought intact during longer brainstorms."
What This Means
Google is positioning Gemini 3.1 Flash Live as a direct performance upgrade for its voice conversation products. The focus on acoustic nuance recognition and background noise filtering suggests competition with other voice-first AI interfaces. The 90+ language support and global rollout across Search Live indicate Google's strategy to make voice interaction a primary interface for search globally. However, specific benchmark data comparing 3.1 Flash Live to competing audio models (OpenAI's real-time API, for example) is not provided.
Related Articles
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks
Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Black Forest Labs Launches FLUX 3 Video, Claims It Beats Seedance 2.0 on Elo Rankings
Black Forest Labs has made FLUX 3 Video generally available via its API, offering up to 20-second HD/Full HD clips with native audio and lip-sync in 14+ languages. The company claims its internal Elo benchmarks put the model ahead of Seedance 2.0, Gemini Omni Flash, and Minimax H3.
Comments
Loading...