model release

Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler

TL;DR

Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.

3 min read
0

Google today launched Gemini 3.5 Transcribe, a new speech-to-text model the company calls its "most precise speech-to-text model yet." The model is already live in Gboard Rambler on Android, the Gemini app for macOS, and Google Antigravity, with Chrome browser support coming next.

Benchmark numbers

According to Google, measurements from Artificial Analysis put Gemini 3.5 Transcribe's average Word Error Rate (WER) at 4.0% for streaming use cases and 2.6% for non-streaming use cases. On the FLEURS multilingual benchmark, the model scores 5.50% WER in streaming mode and 5.04% WER in non-streaming mode across a set of top languages and locales.

Google claims the model improves time to final transcription by 70% compared to its 2025-era Chirp 3 transcription model, alongside better word error rates. These are Google's own reported figures rather than independently verified third-party results, though Artificial Analysis is cited as the measurement source.

What the model does differently

Unlike conventional automatic speech recognition systems, Gemini 3.5 Transcribe converts raw audio directly into polished, formatted text rather than a raw transcript requiring separate cleanup. The model handles self-corrections in speech (e.g., "let's meet Tuesday — no, Wednesday"), strips filler words like "um" and "ah," and auto-formats the resulting text.

Other capabilities Google is touting:

  • Custom vocabulary: adapts transcription to user-supplied specialized jargon and unique spellings.
  • Language coverage: automatically detects and transcribes over 85 languages, including regional accents and dialects.
  • Multi-speaker identification: attributes speech to up to three speakers with timestamps in pre-recorded audio; support beyond three speakers is labeled experimental.
  • Function calling: can delegate tasks like image generation or file analysis to other Gemini models, demonstrated through the "Speak to Window" feature in the Gemini app for macOS.
  • Noisy-environment accuracy: Google says the model accurately captures alphanumeric strings such as postal codes and order IDs in real-world noisy conditions.

In Google Antigravity, the model reportedly pairs screen context and chat history — with user permission — to improve transcription accuracy for file names, agent outputs, and active documents.

Availability

Gemini 3.5 Transcribe is rolling out across several surfaces:

  • Gboard Rambler on Android (live now)
  • Gemini app for macOS, via the Speak to Window feature (live now)
  • Google Antigravity (live now)
  • Chrome browser, for voice dictation into any web field (coming soon)
  • Developers: public preview through the Gemini API via Google AI Studio and Google Antigravity
  • Enterprises: public preview via the Gemini Enterprise Agent Platform, with support for Gemini Enterprise for Customer Experience coming soon

Google has not disclosed pricing for API access to Gemini 3.5 Transcribe, and no context window figure applies in the traditional sense since this is an audio-to-text model rather than a general-purpose language model.

What this means

Gemini 3.5 Transcribe is Google's answer to the growing demand for voice-first interfaces across its product line, from keyboards to browsers to coding tools. The reported WER improvements and 70% latency reduction over Chirp 3 suggest meaningful engineering progress, though the figures come from Google-commissioned benchmarking via Artificial Analysis rather than independent verification. The more significant strategic move may be the function-calling integration — letting voice input trigger downstream Gemini actions like image generation — which points toward voice becoming a control layer for Google's broader agentic tooling, not just a dictation feature. Chrome's upcoming rollout will be the real test of scale, given its user base dwarfs Gboard Rambler or the macOS Gemini app.

Related Articles

model release

Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech

Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.

model release

Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing

Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.

model release

Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Model Matches Larger Rivals at $0.16/$0.47 per Million Tokens

Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates just 6 billion parameters per token and previews architecture planned for Qwen4. The model outperforms the much larger Qwen3.7-Plus at roughly one-ninth the training cost and ships at $0.16 per million input tokens and $0.47 per million output tokens.

model release

Zhipu AI Releases GLM-5.3-Flash: First Multimodal Model in GLM-5 Series, 320B Parameters with Only 18B Active

Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 across benchmarks at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic tasks.

Comments

Loading...