Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing
Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.
What happened
Google DeepMind released Gemini 3.5 Transcribe, a new speech-to-text model built for real-time transcription, on August 26, 2026. The model is available in public preview through the Gemini API in Google AI Studio, the Gemini Enterprise Agent Platform, and Google Antigravity.
The model ships as two separate API endpoints. gemini-3.5-transcribe-live handles continuous, bidirectional streaming with sub-second latency for interactive voice applications via the Live API. gemini-3.5-transcribe processes pre-recorded audio — meetings, call logs, and similar files — with speaker attribution and word-level timestamps via the Interactions API.
Benchmark claims
According to Google, third-party benchmarking firm Artificial Analysis measured an average word error rate (WER) of 4.0% for streaming transcription and 2.6% for non-streaming transcription. On the FLEURS multilingual benchmark, Google reports 5.50% WER in streaming mode and 5.04% WER in non-streaming mode across a set of top languages and locales. Google also claims a 70% improvement in time-to-final-transcription compared to its previous transcription model, Chirp 3, as measured by Artificial Analysis. These figures have not been independently verified by this publication.
The model claims support for over 85 languages with automatic detection, regional accent handling, and dialect recognition. It attributes speech to up to three speakers in pre-recorded audio with timestamps; support beyond three speakers is labeled experimental by Google.
Capabilities
Beyond raw transcription, Gemini 3.5 Transcribe performs what Google calls "smart transcription": cleaning up self-corrections (e.g., "let's meet Tuesday — no, Wednesday"), removing filler words, and auto-formatting output. It also supports function calling, letting the model delegate tasks like image generation or file analysis to other Gemini models — currently live in the Gemini macOS app. The model also accepts custom vocabulary lists for specialized jargon or unusual spellings.
Where it's showing up
Google has already deployed the model in consumer products: Rambler, a dictation feature on Android's Gboard, and the Gemini app on macOS, where it powers voice commands paired with on-screen context. Google Antigravity uses the model with screen context and chat history (with user permission) to improve transcription accuracy for file names and technical terms. Support is coming soon to Chrome for voice input in any web text field. Enterprise availability is coming soon to Gemini Enterprise for Customer Experience.
Third-party platforms including LiveKit, LangChain, Pipecat, Vercel, Agora, Fishjam, and Vision Agents have integrated the Live API to build voice interfaces on top of the model. Google cited Vivo, Intellitek Health, and Lingopal as early enterprise users providing positive feedback on latency, accuracy, and language coverage.
Pricing for either API endpoint has not yet been disclosed.
What this means
This is a product-line update to Google's audio stack rather than a foundation model with broad general reasoning claims — the news is entirely about transcription quality and integration surface area, not about a new large language model. The competitive stakes are in the WER numbers and latency, where Google is directly positioning against Whisper-class models and enterprise transcription vendors like Deepgram and AssemblyAI. The function-calling feature — letting a transcription model trigger downstream Gemini actions — signals Google's intent to make voice input into an app the primary agent front-end, not just a text-input convenience feature. Whether the claimed WER figures hold up outside Google's own benchmarking partner is the detail worth watching once independent evaluations arrive.
Related Articles
Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler
Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.
Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech
Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.
Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Model Matches Larger Rivals at $0.16/$0.47 per Million Tokens
Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates just 6 billion parameters per token and previews architecture planned for Qwen4. The model outperforms the much larger Qwen3.7-Plus at roughly one-ninth the training cost and ships at $0.16 per million input tokens and $0.47 per million output tokens.
Zhipu AI Releases GLM-5.3-Flash: First Multimodal Model in GLM-5 Series, 320B Parameters with Only 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 across benchmarks at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic tasks.
Comments
Loading...