Mistral releases Voxtral TTS, open-source speech model for enterprise voice agents
Mistral AI released Voxtral TTS, an open-source text-to-speech model designed for enterprise voice agents and edge devices. The model supports nine languages, adapts custom voices from samples under five seconds, and achieves 90ms time-to-first-audio latency with a 6x real-time factor.
Mistral AI released Voxtral TTS on Thursday, an open-source text-to-speech model targeting enterprise voice applications and edge deployment. The model directly competes with ElevenLabs, Deepgram, and OpenAI's voice offerings.
Model Specifications
Voxtral TTS supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic. The model is based on Ministral 3B and designed for real-time performance with a time-to-first-audio (TTFA) of 90 milliseconds for a 10-second, 500-character sample. Its real-time factor (RTF) is 6x, meaning it can render a 10-second audio clip in approximately 1.6 seconds.
The model adapts to custom voices from samples shorter than five seconds while preserving accent, inflection, intonation, and speech irregularities. According to Mistral, it can switch between languages without losing voice characteristics—useful for dubbing and real-time translation applications.
Positioning and Capabilities
Pierre Stock, VP of science operations at Mistral AI, told TechCrunch that the company built "a small-sized speech model that can fit on a smartwatch, a smartphone, a laptop, or other edge devices" with "a cost that is a fraction of anything else on the market." The company emphasizes human-sounding output and real-time performance as core differentiators.
Mistral positions the open-source nature and customization flexibility as competitive advantages, allowing enterprises to tune models for specific use cases rather than relying on proprietary, managed solutions.
Strategic Context
Voxtral TTS complements Mistral's earlier 2026 releases of transcription models for batch and real-time processing. Stock indicated the company plans "an end-to-end platform that can handle multimodal streams of input, including audio, text, and image and output as well," suggesting a broader vision for agentic systems that process multiple modalities.
Pricing details were not disclosed. Availability for open-source use or commercial deployment terms remain unspecified.
What this means
Mistral is building a complete voice AI stack to compete with specialized speech companies and large language model providers offering voice capabilities. The open-source release strategy trades proprietary advantage for developer adoption and enterprise customization flexibility. The 90ms latency and edge-device focus suggest targeting real-time conversational agents rather than pre-rendered content, positioning against both traditional TTS vendors and API-based competitors.
Related Articles
Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens
Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Releases MiMo-V2.6-Flash-RL, a 309B-Parameter MoE Model with 1M-Token Context and Native Omnimodal Support
Xiaomi's MiMo team released MiMo-V2.6-Flash-RL, an efficiency-tier checkpoint in the MiMo-V2.6 series featuring a 309B-parameter (15B active) Mixture-of-Experts architecture, 1M-token context, and native support for text, image, video, and audio. The model uses a single mixed reinforcement learning run across coding, agentic, visual, and cybersecurity tasks rather than domain-specific training.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Comments
Loading...