model release

Google Launches Gemini 3.8 Flash TTS: Voice Cloning and Text-Described Voices for $9-18 per Million Audio Tokens

TL;DR

Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, two speech generation models that let users design voices from text descriptions or clone a voice from a 30-second sample. Both support over 100 languages and roll out now through the Gemini API and Google AI Studio.

3 min read
0

Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, that let users design synthetic voices from text descriptions rather than choosing only from a fixed library. Both models support more than 100 languages and are rolling out now through the Gemini API and Google AI Studio.

Text-described voices and cloning

With Flash TTS, a text prompt can define a voice's role, accent, and vocal traits, according to Google. Users who prefer a preset can pick from a library of more than 2,000 built-in voices, including regional variants like Mexican Spanish, Quebec French, and Scottish English.

A voice cloning feature builds a voice profile from a 30-second audio sample. Google requires the person being cloned to record a spoken consent statement that matches the voice in the sample. Every generated clip carries an inaudible SynthID watermark intended to help detect AI-generated speech, according to the company.

Google also previewed a feature called "Voice Remixing," which will let users adjust the timbre, pitch, tempo, and accent of library voices. It is not yet available.

Stage directions and dialogue control

Both models accept per-line stage directions or can interpret script cues automatically. Google says the models can generate hours of audio with minimal "speaker drift" — meaning the voice stays consistent over long outputs. A two-voice mode generates dialogue from a single script while keeping each speaker distinct, and users can script nonverbal sounds like laughter, sighs, and filler words such as "mhm" to control timing and pacing.

Google claims its models lead in most categories of Hume AI's text-to-speech benchmark, though the company has not published specific scores alongside the release.

Two models, two use cases

Gemini 3.8 Flash TTS targets creative production — podcasts, audiobooks, and game character voices. Gemini 3.8 Flash-Lite TTS is positioned for lower-cost speech generation at scale, aimed at dubbing, audio content, and voice agents.

Pricing is billed in US dollars per million tokens, with text tokens for input and audio tokens for output. Through the end of 2026, Flash TTS costs $0.50 per 1M input tokens and $9.00 per 1M output tokens; Flash-Lite TTS costs $0.50 per 1M input tokens and $6.00 per 1M output tokens. Both rates roughly double starting January 1, 2027, rising to $1.00/$18.00 for Flash TTS and $1.00/$12.00 for Flash-Lite TTS.

Google says one second of generated audio equals 25 audio tokens, putting an hour of output at 90,000 tokens. That works out to $0.81 per hour of audio with Flash TTS and $0.54 per hour with Flash-Lite TTS under 2026 pricing, rising to $1.62 and $1.08 respectively in 2027.

Flash TTS is available in Gemini Notebook, and Flash-Lite TTS is available in Google Vids. Gemini Enterprise API access is expected to follow for both models. Developer platforms including Agora, LiveKit, Pipecat, and Vercel already support integration through the Gemini API. Google has not yet listed regional endpoints for the new models, though earlier TTS models offered EU data processing. Data from free-tier usage is used to improve Google's products; paid-tier data is not, according to the company.

What this means

The split between a premium creative model and a low-cost scaled model signals Google is targeting two distinct buyer types: studios producing narrative audio and enterprises running voice agents or dubbing pipelines at volume. Text-described voice design and 30-second cloning lower the barrier to custom voice production, but the consent-recording requirement for cloning shows Google is trying to preempt misuse concerns before regulators do. The near-doubling of prices in 2027 suggests introductory rates meant to drive adoption before Google normalizes margins on the audio output side, which is priced far higher than text input.

Related Articles

model release

Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS with Voice Creation from Text Prompts

Google DeepMind has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models that generate custom voices from natural language prompts and support line-by-line performance direction. The models top Hume AI's Voice Design Benchmark at 71.4 and claim first and second place on its Overall Quality Index.

model release

Alibaba Launches Qwen-Audio-3.1, Cuts AI Audio API Prices by Up to 95 Percent

Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model lineup covering speech recognition, text-to-speech, and real-time voice interaction. Alongside the release, Alibaba cut API pricing by up to 95 percent for ASR, 85 percent for real-time models, and 70 percent for TTS.

model release

Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview

A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.

model release

Aion Labs Launches Aion 3.5 Mini, a $0.70/M-Token Roleplaying Model with 262K Context

Aion Labs has released Aion 3.5 Mini, a lower-cost version of its multi-model roleplaying system Aion 3.5. Built on the GLM model family, it offers a 262K token context window at $0.70 per 1M input tokens and $1.40 per 1M output tokens.

Comments

Loading...