Mistral releases Voxtral-4B-TTS-2603, open-weights text-to-speech model for production voice agents
Mistral AI released Voxtral-4B-TTS-2603, an open-weights text-to-speech model designed for production voice agents. The 4B-parameter model supports 9 languages, 20 preset voices, achieves 70ms latency at concurrency 1 on a single NVIDIA H200, and requires only 16GB GPU memory.
Mistral Releases Voxtral-4B-TTS-2603 Open Text-to-Speech Model
Mistral AI released Voxtral-4B-TTS-2603, an open-weights text-to-speech model built for production voice agent deployment. The model is distributed under CC BY-NC 4 license with BF16 weights and 20 reference voices.
Performance and Hardware Requirements
Voxtral-4B requires a minimum of 16GB GPU memory and runs on a single NVIDIA H200. Measured on vLLM v0.18.0 with 500-character text input and 10-second audio reference:
- Single concurrent request: 70ms latency, 0.103 real-time factor (RTF), 119.14 characters/second/GPU throughput
- 16 concurrent requests: 331ms latency, 0.237 RTF, 879.11 characters/second/GPU throughput
- 32 concurrent requests: 552ms latency, 0.302 RTF, 1,430.78 characters/second/GPU throughput
Language and Voice Support
The model supports 9 languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Arabic, and Hindi. It includes 20 preset voices with dialect diversity and delivers 24kHz audio output in multiple formats (WAV, PCM, FLAC, MP3, AAC, Opus). Voice customization is available through Mistral's AI Studio.
Technical Architecture
Voxtral-4B is fine-tuned from Mistral's Ministral-3-3B-Base-2512 model. The release includes production-grade support through vLLM-Omni (version >= 0.18.0), developed in collaboration with the vLLM team. The model supports streaming and batch inference modes.
Deployment and Licensing
The model ships with vLLM-Omni integration and includes a Docker image option for containerized deployment. Installation requires vllm >= 0.18.0 and mistral_common >= 1.10.0.
The reference voices inherit CC BY-NC 4 licensing from source datasets (EARS, CML-TTS, IndicVoices-R, Arabic Natural Audio). Mistral specifies users must comply with applicable laws and are responsible for avoiding misuse.
Stated Use Cases
Mistral positions Voxtral-4B for customer support, financial services KYC workflows, manufacturing operations, government services, supply chain logistics, in-vehicle systems, sales and marketing, and real-time translation.
What This Means
Voxtral-4B represents Mistral's entry into the open-source TTS space, competing against closed commercial solutions. The sub-100ms latency and 4B parameter count target production deployments with moderate hardware requirements. CC BY-NC licensing restricts commercial use to Mistral's terms, limiting adoption for commercial SaaS applications compared to permissive open licenses. The model's performance at 32 concurrent requests (1,430 characters/second throughput) positions it for real-time voice agent infrastructure, though practical throughput will depend on actual workload patterns and hardware availability.
Related Articles
Google Launches Gemini 3.8 Flash TTS: Voice Cloning and Text-Described Voices for $9-18 per Million Audio Tokens
Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, two speech generation models that let users design voices from text descriptions or clone a voice from a 30-second sample. Both support over 100 languages and roll out now through the Gemini API and Google AI Studio.
Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS with Voice Creation from Text Prompts
Google DeepMind has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models that generate custom voices from natural language prompts and support line-by-line performance direction. The models top Hume AI's Voice Design Benchmark at 71.4 and claim first and second place on its Overall Quality Index.
Alibaba Launches Qwen-Audio-3.1, Cuts AI Audio API Prices by Up to 95 Percent
Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model lineup covering speech recognition, text-to-speech, and real-time voice interaction. Alongside the release, Alibaba cut API pricing by up to 95 percent for ASR, 85 percent for real-time models, and 70 percent for TTS.
Nvidia Releases Nemotron 3 Diarization, a Free 100M-Parameter Model That Tracks 8 Speakers in Real Time
Nvidia released Nemotron 3 Diarization, a free 100-million-parameter model that identifies who is speaking in real time across up to eight participants. It leads the VoiceArena Diarization Benchmark v1 with a 14.7% error rate, cutting errors by 41% versus its predecessor.
Comments
Loading...