Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS with Voice Creation from Text Prompts
Google DeepMind has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models that generate custom voices from natural language prompts and support line-by-line performance direction. The models top Hume AI's Voice Design Benchmark at 71.4 and claim first and second place on its Overall Quality Index.
Google DeepMind released two new text-to-speech models on September 23, 2026: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The models are available now in Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Gemini 3.8 Flash TTS is positioned for creative direction and character design, letting users generate original voices from natural-language prompts across more than 100 languages and dialects. Gemini 3.8 Flash-Lite TTS targets high-volume use cases like dubbing and voice agents, optimized for cost efficiency at scale. Pricing not yet disclosed for either model.
Benchmark Claims
According to Google, Gemini 3.8 Flash TTS holds the #1 overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and leads in accent modeling with a score of 60.8. The company also claims Gemini 3.8 Flash TTS and Flash-Lite TTS take the #1 and #2 spots, respectively, on Hume AI's Overall Quality Index. In blind human preference evaluations on Voice Arena, Google states the models rank at the top among competitors for Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi. None of these benchmark claims have been independently verified.
What's New
The models expand voice options from 30 preset voices to what Google describes as an "infinite library" through generative voice design — users describe a voice ("a fire-breathing dragon" or "a charismatic narrator with a regional cadence") and the model generates it. Google AI Studio also ships with over 2,000 production-ready voices covering regional variants like Mexican Spanish, Quebec French, and Scots English.
Voice replication lets users recreate a vocal profile from a 30-second audio sample, gated by a consent-verification step requiring a matching verbal consent recording from the voice owner. Other features include line-by-line performance direction with acting cues, long-form generation aimed at minimizing speaker drift across hours of audio, native two-speaker scene staging for dialogue, and scripted non-verbal cues like laughs, sighs, and backchanneling interjections.
A "voice remixing" feature — letting users adjust timbre, pitch, pace, and accent of existing library voices via prompts — is listed as coming soon rather than available at launch.
Every generated audio clip is watermarked with SynthID, and outputs carry C2PA credentials for provenance tracking. Google says this is intended to keep AI-generated speech detectable and reduce misinformation risk.
The release follows other entries in Google's Gemini Audio lineup, including 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking, and is compared favorably by Google against its predecessor, Gemini 3.1 Flash TTS, particularly on long-form content and dual-speaker screenplay control.
What this means
Google is pushing TTS beyond fixed-voice presets toward a prompt-driven voice design workflow, competing directly with players like ElevenLabs in synthetic voice generation. The consent-verification requirement for voice cloning and mandatory SynthID watermarking signal Google is trying to get ahead of misuse concerns before regulators force the issue. The lack of published pricing makes it hard to judge whether Flash-Lite TTS actually delivers on its cost-efficiency promise for high-volume applications like dubbing and voice agents — that comparison will matter more than benchmark rankings once developers start building at scale.
Related Articles
Alibaba Launches Qwen-Audio-3.1, Cuts AI Audio API Prices by Up to 95 Percent
Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model lineup covering speech recognition, text-to-speech, and real-time voice interaction. Alongside the release, Alibaba cut API pricing by up to 95 percent for ASR, 85 percent for real-time models, and 70 percent for TTS.
NVIDIA Releases Nemotron 3 Diarization, a 100M-Parameter Open-Weight Model Ranked #1 on Voice Arena's Diarization-Bench
NVIDIA released Nemotron 3 Diarization, a 100-million-parameter open-weight model that identifies who is speaking and when in audio conversations. It ranked #1 among 17 system configurations on Voice Arena's Diarization-Bench with a 14.72% diarization error rate, supporting up to eight speakers in both live and recorded audio.
Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview
A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.
Aion Labs Launches Aion 3.5 Mini, a $0.70/M-Token Roleplaying Model with 262K Context
Aion Labs has released Aion 3.5 Mini, a lower-cost version of its multi-model roleplaying system Aion 3.5. Built on the GLM model family, it offers a 262K token context window at $0.70 per 1M input tokens and $1.40 per 1M output tokens.
Comments
Loading...