model release

Meta Launches Muse Voice Transcribe, Real-Time Speech Model Handling 20+ Speakers Across 70+ Languages

TL;DR

Meta Superintelligence Lab released Muse Voice Transcribe, its first real-time audio perception model, claiming state-of-the-art streaming speech-to-text with native speaker diarization. The model handles 20+ speakers and code-switching across languages, priced at $3 per 1,000 audio minutes.

3 min read
0

Meta Superintelligence Lab (MSL) released Muse Voice Transcribe on September 1, 2026, a real-time audio perception model that Meta claims can distinguish between more than 20 speakers and switch between languages mid-sentence.

The model performs speaker diarization and endpointing natively in a single system — a departure from pipelines that typically chain separate models for these tasks together. According to Meta CEO Mark Zuckerberg, who demonstrated the model on X, Muse Voice Transcribe uses "adaptive delay to predict each token," waiting longer on ambiguous words and committing faster on straightforward ones to balance latency against accuracy.

Training and language coverage

Meta says the model was trained across more than 70 languages, with 25 validated at launch. It claims the system handles "code-switching" — sentences that blend vocabulary from multiple languages mid-utterance — and can sustain accuracy across hour-long sessions involving 20 or more distinct speakers. Meta has not published independent benchmark scores or a technical report alongside the release, so these performance claims remain unverified by third parties.

Availability and pricing

Muse Voice Transcribe is live now through three channels: it already powers dictation inside Meta's recently launched Meta AI Mac app, it's integrated into Muse Code (Meta's coding agent), and it's accessible to developers via Meta's Model API. Pricing is set at $3 per 1,000 audio minutes. A demo version is also available on Meta's research blog for anyone to test without an API key.

Meta AI chief scientist Alexandr Wang confirmed on X that the model is "live now via meta model API," already powering dictation across the desktop app and coding tool.

Competitive context

The release lands less than a week after Google shipped Gemini 3.5 Transcribe, an audio model with comparable multi-speaker and multilingual capabilities. Google has said it plans to build its model into Android and eventually Chrome. Meta has not disclosed similar plans to embed Muse Voice Transcribe into WhatsApp, Instagram, or its other consumer-facing products — for now, the model's visible reach is limited to the Mac app and developer API.

Muse Voice Transcribe is the latest in a rapid string of releases from MSL, which has shipped a dedicated coding agent, an open-weight language model, and the Meta AI Mac app all within the past few weeks.

What this means

Meta is positioning MSL as a high-velocity model shop, shipping specialized systems (audio, code, open-weight) in quick succession rather than a single flagship release cycle. The $3-per-1,000-minutes price point undercuts many enterprise transcription APIs, and native diarization removes a step most speech-to-text pipelines currently handle with bolt-on models — a real technical convenience if the claimed accuracy holds up under independent testing. But the near-simultaneous launch with Google's Gemini 3.5 Transcribe suggests real-time multilingual, multi-speaker transcription is becoming a contested commodity feature rather than a differentiator. The bigger question is distribution: Google is wiring its model into Android's operating system layer, while Meta's plans for its billion-plus-user apps (WhatsApp, Instagram) remain unstated. Without that integration, Muse Voice Transcribe risks staying a developer-tier product rather than a consumer-facing capability.

Related Articles

model release

Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech

Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.

model release

Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages

Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.

model release

Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing

Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.

model release

Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler

Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.

Comments

Loading...