Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler
Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.
Google today launched Gemini 3.5 Transcribe, a new speech-to-text model the company calls its "most precise speech-to-text model yet." The model is already live in Gboard Rambler on Android, the Gemini app for macOS, and Google Antigravity, with Chrome browser support coming next.
Benchmark numbers
According to Google, measurements from Artificial Analysis put Gemini 3.5 Transcribe's average Word Error Rate (WER) at 4.0% for streaming use cases and 2.6% for non-streaming use cases. On the FLEURS multilingual benchmark, the model scores 5.50% WER in streaming mode and 5.04% WER in non-streaming mode across a set of top languages and locales.
Google claims the model improves time to final transcription by 70% compared to its 2025-era Chirp 3 transcription model, alongside better word error rates. These are Google's own reported figures rather than independently verified third-party results, though Artificial Analysis is cited as the measurement source.
What the model does differently
Unlike conventional automatic speech recognition systems, Gemini 3.5 Transcribe converts raw audio directly into polished, formatted text rather than a raw transcript requiring separate cleanup. The model handles self-corrections in speech (e.g., "let's meet Tuesday — no, Wednesday"), strips filler words like "um" and "ah," and auto-formats the resulting text.
Other capabilities Google is touting:
- Custom vocabulary: adapts transcription to user-supplied specialized jargon and unique spellings.
- Language coverage: automatically detects and transcribes over 85 languages, including regional accents and dialects.
- Multi-speaker identification: attributes speech to up to three speakers with timestamps in pre-recorded audio; support beyond three speakers is labeled experimental.
- Function calling: can delegate tasks like image generation or file analysis to other Gemini models, demonstrated through the "Speak to Window" feature in the Gemini app for macOS.
- Noisy-environment accuracy: Google says the model accurately captures alphanumeric strings such as postal codes and order IDs in real-world noisy conditions.
In Google Antigravity, the model reportedly pairs screen context and chat history — with user permission — to improve transcription accuracy for file names, agent outputs, and active documents.
Availability
Gemini 3.5 Transcribe is rolling out across several surfaces:
- Gboard Rambler on Android (live now)
- Gemini app for macOS, via the Speak to Window feature (live now)
- Google Antigravity (live now)
- Chrome browser, for voice dictation into any web field (coming soon)
- Developers: public preview through the Gemini API via Google AI Studio and Google Antigravity
- Enterprises: public preview via the Gemini Enterprise Agent Platform, with support for Gemini Enterprise for Customer Experience coming soon
Google has not disclosed pricing for API access to Gemini 3.5 Transcribe, and no context window figure applies in the traditional sense since this is an audio-to-text model rather than a general-purpose language model.
What this means
Gemini 3.5 Transcribe is Google's answer to the growing demand for voice-first interfaces across its product line, from keyboards to browsers to coding tools. The reported WER improvements and 70% latency reduction over Chirp 3 suggest meaningful engineering progress, though the figures come from Google-commissioned benchmarking via Artificial Analysis rather than independent verification. The more significant strategic move may be the function-calling integration — letting voice input trigger downstream Gemini actions like image generation — which points toward voice becoming a control layer for Google's broader agentic tooling, not just a dictation feature. Chrome's upcoming rollout will be the real test of scale, given its user base dwarfs Gboard Rambler or the macOS Gemini app.
Related Articles
Google Lists Gemini 3.8 Flash on OpenRouter With 1M-Token Context, September 2026 Release Date
Google's Gemini 3.8 Flash has surfaced on OpenRouter with a 1-million-token context window and discounted pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens. Google has not issued a separate public announcement, and the listed release date of September 2, 2026 is unusually far out, leaving key details unconfirmed.
Google Launches WeatherNext 3, Claims 50% More Accurate Precipitation Forecasts
Google DeepMind and Google Research released WeatherNext 3, a weather AI model trained on real-time geostationary satellite data instead of lagging numerical weather prediction outputs. Google claims up to 50% more accurate day-ahead precipitation forecasts, now rolling out to Search, Maps, and the Gemini app.
Google Launches Gemini 3.8 Flash With Same Pricing as Predecessor, Its Third Flash Model in Six Weeks
Google released Gemini 3.8 Flash, its third Flash model in six weeks, pricing it identically to its predecessor at $0.75 per million input tokens and $3.75 per million output tokens. The launch coincides with a favorable antitrust ruling and public praise from Berkshire Hathaway's Greg Abel, giving Alphabet a stronger narrative after a four-month stock slide.
OpenAI's GPT-6 Astra Cuts Hallucinations, But Indirect Prompt Injection Attacks Still Succeed 8.5% of the Time
OpenAI's new GPT-6 Astra model shows major improvements in hallucination rates and jailbreak resistance over predecessor GPT-5.6 Sol, according to OpenAI's system card. However, indirect prompt injection attacks hidden in documents still succeed 8.5% of the time in external testing by Gray Swan, down from 27% but still above rival Claude Opus 5's 4.8% rate.
Comments
Loading...