Mistral releases Voxtral, open-weight TTS model that clones voices from 3 seconds of audio
Mistral has released Voxtral TTS, a 4-billion-parameter text-to-speech model that can clone voices from just three seconds of reference audio across nine languages. The model delivers 70ms latency for typical 10-second samples and outperformed ElevenLabs Flash v2.5 in naturalness tests. Voxtral is available via API at $0.016 per 1,000 characters and as open-weights on Hugging Face.
Mistral Releases Voxtral: Open-Weight TTS Model with Voice Cloning from 3-Second Samples
Mistral has released Voxtral TTS, its first text-to-speech model, positioning it as a compact alternative to closed proprietary systems. The model contains 4 billion parameters and supports nine languages: German, English, French, Spanish, and five others.
Key Technical Specifications
Voxtral's standout capability is voice cloning from minimal audio. The model requires just three seconds of reference audio to adapt to and replicate new voices, with support for emotionally expressive speech synthesis. Latency benchmarks show 70 milliseconds for a typical configuration processing 10-second speech samples with 500 characters of input text.
The model operates across a broader linguistic range than many competing TTS systems, though Mistral has not specified the complete language list beyond the four named examples.
Performance vs. Competitors
In human evaluation tests, Voxtral TTS scored higher on naturalness compared to ElevenLabs Flash v2.5 at comparable response times. However, this comparison has a timing caveat: ElevenLabs subsequently released version 3, which was not included in Mistral's evaluation. This means the benchmark reflects performance against a prior-generation ElevenLabs model rather than current-generation alternatives.
Availability and Pricing
Mistral offers three access paths for Voxtral TTS:
- API access: $0.016 per 1,000 characters
- Mistral Studio: Web-based testing interface
- Open-weights version: Available on Hugging Face for local deployment and fine-tuning
The open-weights release represents a departure from Mistral's approach with some of its larger language models, giving developers the ability to run Voxtral locally without relying on the company's infrastructure.
What This Means
Voxtral establishes Mistral as a competitor in the TTS market beyond its core language modeling business. The 4-billion-parameter size makes it accessible for resource-constrained deployments—substantially smaller than many alternatives—while the open-weights availability appeals to enterprises avoiding vendor lock-in. The three-second voice cloning threshold is practically significant, reducing friction for users who need quick voice adaptation. The API pricing at $0.016 per 1,000 characters is competitive but not a market undercut; comparison requires converting to per-token equivalents based on language-specific tokenization rates. The main strategic value lies in the open-source option, which appeals to builders wanting fine-tuning and deployment flexibility that proprietary APIs don't provide.
Related Articles
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens
Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.
Anthropic Ships Claude Opus 5.5, OpenAI Counters with GPT-6 Sol and Luna Hours Later, Triggering Sharp Price Cuts
Anthropic released Claude Opus 5.5 with a 20% price cut, and roughly an hour later OpenAI shipped GPT-6 Sol and GPT-6 Luna at roughly half the price of their GPT-5.6 predecessors. The releases follow Grok 4.7 and MiMo v2.6 from the previous day, intensifying competition among frontier model providers.
Comments
Loading...