model release

Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler

TL;DR

Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.

3 min read
0

Google today launched Gemini 3.5 Transcribe, a new speech-to-text model the company calls its "most precise speech-to-text model yet." The model is already live in Gboard Rambler on Android, the Gemini app for macOS, and Google Antigravity, with Chrome browser support coming next.

Benchmark numbers

According to Google, measurements from Artificial Analysis put Gemini 3.5 Transcribe's average Word Error Rate (WER) at 4.0% for streaming use cases and 2.6% for non-streaming use cases. On the FLEURS multilingual benchmark, the model scores 5.50% WER in streaming mode and 5.04% WER in non-streaming mode across a set of top languages and locales.

Google claims the model improves time to final transcription by 70% compared to its 2025-era Chirp 3 transcription model, alongside better word error rates. These are Google's own reported figures rather than independently verified third-party results, though Artificial Analysis is cited as the measurement source.

What the model does differently

Unlike conventional automatic speech recognition systems, Gemini 3.5 Transcribe converts raw audio directly into polished, formatted text rather than a raw transcript requiring separate cleanup. The model handles self-corrections in speech (e.g., "let's meet Tuesday — no, Wednesday"), strips filler words like "um" and "ah," and auto-formats the resulting text.

Other capabilities Google is touting:

  • Custom vocabulary: adapts transcription to user-supplied specialized jargon and unique spellings.
  • Language coverage: automatically detects and transcribes over 85 languages, including regional accents and dialects.
  • Multi-speaker identification: attributes speech to up to three speakers with timestamps in pre-recorded audio; support beyond three speakers is labeled experimental.
  • Function calling: can delegate tasks like image generation or file analysis to other Gemini models, demonstrated through the "Speak to Window" feature in the Gemini app for macOS.
  • Noisy-environment accuracy: Google says the model accurately captures alphanumeric strings such as postal codes and order IDs in real-world noisy conditions.

In Google Antigravity, the model reportedly pairs screen context and chat history — with user permission — to improve transcription accuracy for file names, agent outputs, and active documents.

Availability

Gemini 3.5 Transcribe is rolling out across several surfaces:

  • Gboard Rambler on Android (live now)
  • Gemini app for macOS, via the Speak to Window feature (live now)
  • Google Antigravity (live now)
  • Chrome browser, for voice dictation into any web field (coming soon)
  • Developers: public preview through the Gemini API via Google AI Studio and Google Antigravity
  • Enterprises: public preview via the Gemini Enterprise Agent Platform, with support for Gemini Enterprise for Customer Experience coming soon

Google has not disclosed pricing for API access to Gemini 3.5 Transcribe, and no context window figure applies in the traditional sense since this is an audio-to-text model rather than a general-purpose language model.

What this means

Gemini 3.5 Transcribe is Google's answer to the growing demand for voice-first interfaces across its product line, from keyboards to browsers to coding tools. The reported WER improvements and 70% latency reduction over Chirp 3 suggest meaningful engineering progress, though the figures come from Google-commissioned benchmarking via Artificial Analysis rather than independent verification. The more significant strategic move may be the function-calling integration — letting voice input trigger downstream Gemini actions like image generation — which points toward voice becoming a control layer for Google's broader agentic tooling, not just a dictation feature. Chrome's upcoming rollout will be the real test of scale, given its user base dwarfs Gboard Rambler or the macOS Gemini app.

Related Articles

model release

Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context

Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.

model release

OpenAI launches GPT-6 in ChatGPT with 'Intelligent UI' and interactive answers; Sol for paid users, Luna for free

OpenAI is rolling out GPT-6 to all ChatGPT tiers, with paying users on GPT-6 Sol and free users on GPT-6 Luna. The release adds 'Intelligent UI,' which renders answers as interactive charts, buttons, forms and mini apps, and lets the model respond while still thinking. OpenAI claims this cuts wait times by 44 percent.

model release

Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0

Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.

model release

Microsoft releases Decision-1, a Qwen3.5-9B-based model for classification and routing, at $0.042 per 1M input tokens

Microsoft has released Decision-1, a decision model built on Qwen3.5-9B for classification, evaluation, and routing. Microsoft claims 83.5% accuracy across 36 benchmarks and 85 ms latency. Input tokens cost $0.042 per million, and output tokens are free.

Comments

Loading...