Google DeepMind releases Gemini 3.1 Flash TTS with audio tags for precise speech control across 70+ languages
Google DeepMind launched Gemini 3.1 Flash TTS, a text-to-speech model that achieved an Elo score of 1,211 on the Artificial Analysis TTS leaderboard. The model introduces audio tags that allow developers to control vocal style, pace, and delivery through natural language commands embedded in text input, with support for 70+ languages.
Gemini 3.1 Flash TTS — Quick Specs
Google DeepMind releases Gemini 3.1 Flash TTS with audio tags for precise speech control across 70+ languages
Google DeepMind launched Gemini 3.1 Flash TTS, a text-to-speech model that achieved an Elo score of 1,211 on the Artificial Analysis TTS leaderboard based on thousands of blind human preferences. The model is now available in preview via the Gemini API, Google AI Studio, Vertex AI for enterprises, and Google Vids for Workspace users.
Audio tags for granular control
The defining feature of 3.1 Flash TTS is audio tags — natural language commands embedded directly into text input that control vocal style, pace, and delivery. According to Google, these tags provide "improved levels of granularity" for steering AI speech output.
Google AI Studio offers three levels of control:
- Scene direction: Developers can define the environment and provide dialogue instructions to help characters remain "in-character" across multiple turns
- Speaker-level specificity: Unique Audio Profiles can be assigned to characters, with Director's Notes to adjust pace, tone, and accent
- Inline tags: Speakers can change expression mid-sentence, pivoting from high-level settings
Once configured, these parameters can be exported as Gemini API code for consistent voice reproduction across projects.
Performance and availability
Artificial Analysis positioned Gemini 3.1 Flash TTS in its "most attractive quadrant" for combining high-quality speech generation with low cost, though specific pricing was not disclosed. The model supports 70+ languages with what Google describes as "high-fidelity speech" and native multi-speaker dialogue.
All audio generated by the model includes SynthID watermarking — an imperceptible watermark embedded in the audio output designed to enable detection of AI-generated content.
What this means
The introduction of audio tags represents a shift toward programmatic control of AI speech synthesis through natural language rather than complex parameter tuning. The 1,211 Elo score suggests competitive performance against existing TTS models, though direct comparisons to specific competitors weren't provided. The emphasis on multi-speaker dialogue and scene-level control indicates Google is targeting use cases beyond simple text reading — particularly interactive applications, content creation, and localized media production. The mandatory SynthID watermarking addresses growing concerns about audio deepfakes, though the effectiveness of such watermarks against determined adversaries remains an open question in the field.
Related Articles
Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages
Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.
GLM-5.3-Flash Debuts as Zhipu AI's First Natively Multimodal Model, 320B Parameters with 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 while approaching Claude Opus 4.8 on coding and agentic benchmarks at a fraction of the cost. Unsloth has published quantized GGUF versions for local inference.
Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2
Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
Comments
Loading...