Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech
Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.
Google has released Gemini 3.5 Transcribe, a new speech-to-text model the company says can turn unstructured, rambling speech into clean, formatted text. The model joins Gemini 3.5 Live and Gemini 3.5 Live Experimental to round out Google's Gemini Audio family.
According to Google, Gemini 3.5 Transcribe improves on earlier transcription models with greater accuracy and automatic detection of more than 85 languages. The model is designed to handle natural speech patterns rather than requiring users to dictate cleanly — Google claims it can seamlessly process self-corrections, understand a speaker's natural intent, and strip out filler words to produce polished output.
The model also supports voice-based editing commands, letting users revise text they've already dictated using spoken instructions rather than typing. Google says Gemini 3.5 Transcribe can learn custom vocabulary and unique spellings, and is particularly capable at capturing alphanumeric strings like order numbers and postal codes — a detail that points toward enterprise and customer-service use cases.
For multi-speaker audio, Google claims the model can attribute speech to up to three speakers with word-level timestamps when working from pre-recorded audio, a feature suited to transcribing podcasts, meetings, or interviews. No benchmark scores, context window size, or pricing have been disclosed for Gemini 3.5 Transcribe.
The model is already live in two places: it powers the Rambler dictation feature on Android devices, including the Pixel 11 series, and it runs the Gemini app on macOS, where Google says it can work alongside other Gemini models to complete agentic tasks.
Google says broader distribution is coming soon. Chrome users will reportedly be able to use Gemini 3.5 Transcribe for speech-to-text in any web field — dictating replies, social posts, or prompts to Gemini directly from a browser. The model is also available now in Google Antigravity, the company's agentic development platform, and Google says it is coming to Search Live, Gemini Live, Docs, Keep, and Gmail. Developers will get API access to build the model into their own products.
What this means
Gemini 3.5 Transcribe is less about raw transcription accuracy — a race Google, OpenAI, and others have run for years — and more about post-processing: turning messy spoken input into text a person could actually publish or send without editing. That positions it as an assistant-layer feature rather than a standalone product, embedded directly into Android, macOS, Chrome, and Google's productivity suite rather than sold as a separate API-first offering.
The alphanumeric-capture and custom-vocabulary claims suggest Google is targeting business and support workflows, not just casual dictation. But with no benchmark scores or pricing disclosed, independent verification of the model's claimed language coverage and speaker-attribution accuracy will have to wait until developers get broader API access.
Related Articles
Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler
Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.
Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing
Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.
Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Model Matches Larger Rivals at $0.16/$0.47 per Million Tokens
Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates just 6 billion parameters per token and previews architecture planned for Qwen4. The model outperforms the much larger Qwen3.7-Plus at roughly one-ninth the training cost and ships at $0.16 per million input tokens and $0.47 per million output tokens.
Zhipu AI Releases GLM-5.3-Flash: First Multimodal Model in GLM-5 Series, 320B Parameters with Only 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 across benchmarks at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic tasks.
Comments
Loading...