model release

Google DeepMind Releases Gemini 3.5 Live Translate for Real-Time Speech Translation Across 70+ Languages

TL;DR

Google DeepMind released Gemini 3.5 Live Translate, an audio model that provides near real-time speech-to-speech translation across 70+ languages. The model automatically detects languages, preserves speaker intonation and pacing, and maintains a few seconds of latency while generating continuous speech output.

2 min read
0

Google DeepMind Releases Gemini 3.5 Live Translate for Real-Time Speech Translation Across 70+ Languages

Google DeepMind released Gemini 3.5 Live Translate on June 9, 2026, an audio model that provides near real-time speech-to-speech translation across 70+ languages with automatic language detection.

Technical Capabilities

The model generates continuous translated speech while maintaining a latency of "just a few seconds" behind the speaker, according to Google. Unlike turn-based translation systems that wait for complete sentences, Gemini 3.5 Live Translate processes streaming audio and balances translation speed with contextual accuracy.

Key technical features include:

  • Automatic detection of 70+ languages without manual configuration
  • Preservation of speaker intonation, pacing, and pitch in translated output
  • Noise robustness for unpredictable environments
  • Support for over 2,000 language pair combinations in single sessions
  • SynthID watermarking embedded in all generated audio

Availability and Deployment

Gemini 3.5 Live Translate is rolling out across three channels:

Gemini Live API: Available in public preview for developers via Google AI Studio. Developer platforms including Agora, Fishjam, LiveKit, Pipecat, and Vision Agents have integrated the API for real-time media streaming infrastructure.

Google Meet: Launching in private preview this month for select Google Workspace business customers, expanding from the previous limitation of five languages and English-only translation pairs. Broader rollout planned for later in 2026.

Google Translate app: Rolling out globally on Android and iOS. The model powers the Live translate feature for users with connected headphones. Android users receive an additional "listening mode" that streams translations through the phone's earpiece without headphones.

Early Implementations

Grab, which processes over 10 million voice calls monthly, is testing the model to enable multilingual communication between drivers and travelers. Additional partners including CJ ENM and LiveKit have provided feedback on translation quality and low latency, according to Google.

Pricing for API access has not been disclosed.

What This Means

Gemini 3.5 Live Translate represents Google's entry into the competitive real-time speech translation market, directly challenging established players in multilingual communication tools. The 70+ language support and 2,000+ language pair combinations significantly exceed the capabilities of Google's previous Meet translation system, which supported only five languages with English as a required pivot.

The model's continuous streaming approach addresses a core limitation of turn-based systems, though the "few seconds" latency specification lacks precision for developers evaluating real-time requirements. The integration across Google's product ecosystem—from developer APIs to consumer apps—indicates a platform play rather than a standalone model release. However, the lack of disclosed API pricing and benchmark comparisons to competing speech translation models limits technical evaluation.

Related Articles

model release

Google delays Gemini 3.5 Pro release after disappointing coding performance in June training update

Google has delayed the release of Gemini 3.5 Pro past its June deadline due to coding performance issues. The company retrained the model in late June with new data but saw disappointing results, according to Bloomberg. An upgraded Flash model is now in testing with partners.

model release

Moonshot AI releases Kimi K3, China's largest model at 2.8 trillion parameters

Beijing-based Moonshot AI released Kimi K3, China's largest AI model at 2.8 trillion parameters. The company claims the model consistently outperforms OpenAI's GPT 5.5 and Anthropic's Claude Opus 4.8 on benchmarks including coding and general agents, though it still trails the leading-edge GPT 5.6 Sol and Claude Fable 5 in overall performance.

model release

Moonshot AI releases Kimi K3 with 2.7 trillion parameters, claims performance on par with Anthropic Fable 5

Moonshot AI released Kimi K3 on July 16, 2026, featuring 2.7 trillion parameters—the largest open-weight model to date. The company claims K3 performs competitively with Anthropic's Fable 5 while costing $15 per million output tokens compared to Fable's $50.

model release

Moonshot AI's Kimi k3 claims top performance among Chinese models with 1M token context

Moonshot AI has released Kimi k3, positioning it as China's leading AI model. The company claims the model features a 1 million token context window and improved reasoning capabilities, though independent benchmarks are not yet available.

Comments

Loading...