model releaseMicrosoft

Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages

TL;DR

Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.

3 min read
0

Microsoft AI has released MAI-Transcribe-2-Streaming, a real-time speech-to-text model that supports 60 languages and delivers its first partial results in just over 100 milliseconds. Microsoft also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. All three are aimed at voice agents.

MAI-Transcribe-2-Streaming

  • Languages: 60
  • Latency: first partial results in just over 100 ms
  • Price: $0.54 per hour of audio at an introductory rate that runs through the end of the year. Post-promotion pricing has not been disclosed.
  • Accuracy: Microsoft claims the model ranks first for accuracy on Artificial Analysis. No word error rate or other numeric score was provided, and the ranking has not been independently confirmed here.

Microsoft says the low latency lets voice agents begin responding while a user is still mid-sentence.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash

MAI-Voice-2.1 is built to speak 23 languages in the same voice, with a native accent in each one, according to Microsoft.

The Flash variant targets speed and cost:

  • Latency: 150 ms, according to Microsoft
  • Price: $15 per million characters, down from $22. The source does not state which model the $22 figure refers to, but it is presumably the standard MAI-Voice-2.1 rate.

Both voice models can clone a voice from a few seconds of reference audio. Microsoft says built-in safeguards are meant to prevent misuse but has not detailed how they work. In a Microsoft test, about half of 4,000 participants judged the synthetic voices to belong to a real person.

Availability

The models are available through Microsoft Foundry and the MAI Playground, among other platforms. The two voice models are also listed on OpenRouter. The source does not say whether the transcription model is on OpenRouter.

Not disclosed

Parameter counts, training data cutoffs, context limits, and per-model benchmark figures were not published in the announcement coverage. The standard MAI-Voice-2.1 price is also not explicitly stated.

What this means

The release targets the latency budget of conversational voice agents. A first-partial time near 100 ms for transcription and 150 ms for speech synthesis leaves more room for the language model in between. Cascaded pipelines (speech-to-text, LLM, text-to-speech) have usually struggled with that budget compared with end-to-end speech models.

The pricing is aggressive but time-limited. The $0.54 per hour transcription rate is introductory, so teams should wait for post-promotion pricing before committing production volume. The Flash tier's $15 per million characters undercuts the $22 comparison point Microsoft cites.

Listing the voice models on OpenRouter suggests Microsoft wants distribution beyond its own cloud, which would put the MAI line in direct competition with specialist voice vendors such as ElevenLabs. The "ranks first" accuracy claim rests on a single third-party leaderboard and comes without published error rates, so developers should test on their own audio, especially for accented speech, noisy channels, and lower-resource languages among the 60 supported. The voice-cloning capability from seconds of audio, paired with a test where half of listeners took the output for a human, makes the unspecified safeguards the detail to watch.

Related Articles

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

model release

Google unveils Gemini 4 Argon at $2/$10 per 1M tokens, but access is limited to select users

Google unveiled Gemini 4 Argon on Wednesday with introductory pricing of $2 per 1M input tokens and $10 per 1M output tokens, matching OpenAI's discounted GPT-6.1 Sol. Access is restricted to select cybersecurity defenders and enterprise cloud customers, and Google says published rates will double later.

model release

Ideogram 4.5 launches with native 2K output and four tiers from $0.008 to $0.22 per image

Ideogram has released Ideogram 4.5, an image model it claims edits only the area a user specifies and leaves the rest untouched. It offers four quality tiers from 0.8 to 22 cents per image, all at native 2K resolution, via the Ideogram platform and API. An open-weight release is promised but not yet dated.

model release

Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base

Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.

Comments

Loading...