speech-to-text
12 articles tagged with speech-to-text
Meta Launches Muse Voice Transcribe, Real-Time Speech Model Handling 20+ Speakers Across 70+ Languages
Meta Superintelligence Lab released Muse Voice Transcribe, its first real-time audio perception model, claiming state-of-the-art streaming speech-to-text with native speaker diarization. The model handles 20+ speakers and code-switching across languages, priced at $3 per 1,000 audio minutes.
Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages
Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.
Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech
Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.
Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing
Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.
Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler
Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.
OpenAI Releases Whisper Large-v3, Cutting Speech Recognition Errors 10-20% Across Languages
OpenAI has released Whisper large-v3, an open-weight automatic speech recognition and translation model trained on 5 million hours of audio. The model reduces transcription errors by 10-20% compared to its predecessor and adds native Cantonese support.
OpenAI Launches GPT-Live Voice Model That Delegates Complex Tasks to GPT-5.5
OpenAI has replaced ChatGPT's voice mode with GPT-Live, a new voice model that can delegate complex tasks to GPT-5.5 in the background. The previous voice mode was based on a GPT-4o era model with a 2024 knowledge cutoff.
AWS SageMaker AI adds bidirectional streaming for real-time speech transcription with vLLM
Amazon SageMaker AI has launched bidirectional streaming support for real-time inference, enabling WebSocket-based voice applications through vLLM integration. The feature uses HTTP/2 on port 8443 to bridge client connections with vLLM's Realtime API, allowing audio to stream in while transcription streams back simultaneously over a single persistent connection.
OpenAI launches GPT-Realtime-2 with GPT-5-class reasoning, adds real-time translation across 70 languages
OpenAI has added three voice intelligence features to its Realtime API: GPT-Realtime-2 with GPT-5-class reasoning for complex conversational requests, GPT-Realtime-Translate supporting 70 input languages and 13 output languages, and GPT-Realtime-Whisper for live speech-to-text transcription. Translation and transcription are billed by the minute, while GPT-Realtime-2 uses token-based pricing.
Microsoft's MAI-Transcribe-1 achieves lowest word error rate on FLEURS, costs $0.36/audio hour
Microsoft has released MAI-Transcribe-1, a speech-to-text model that achieves the lowest word error rate on the FLEURS benchmark across 25 languages, outperforming Whisper-large-V3, GPT-Transcribe, and Gemini 3.1 Flash-Lite. The model runs 2.5 times faster than Microsoft's previous Azure Fast offering and costs $0.36 per audio hour.
Hume AI open-sources TADA: speech model 5x faster than rivals with zero hallucination
Hume AI has open-sourced TADA, a speech generation model that maps exactly one audio signal to each text token, achieving 5x faster processing than comparable systems. The model produced zero transcription hallucinations across 1,000+ test samples and runs on smartphones, available in 1B and 3B parameter versions under MIT license.
ElevenLabs and Google lead Artificial Analysis speech-to-text benchmark
Artificial Analysis has released an updated speech-to-text benchmark showing ElevenLabs and Google as top performers. The benchmark provides comparative analysis of current speech recognition systems across multiple models.