model release

Google releases Gemini Omni Flash video generation model with conversational editing, withholds speech synthesis

TL;DR

Google DeepMind released Gemini Omni Flash, the first model in its new Omni family that generates and edits video from image, audio, video, and text inputs. The model is rolling out to Gemini app subscribers and YouTube Shorts with a 10-second clip limit, while speech-editing capabilities remain withheld pending safety testing.

2 min read
0

Gemini Omni Flash — Quick Specs

Context window131K tokens
Input$1.5/1M tokens
Output$9/1M tokens

Google releases Gemini Omni Flash video generation model with conversational editing, withholds speech synthesis

Google DeepMind released Gemini Omni Flash on Tuesday at I/O 2026, a multimodal video generation model that accepts any combination of image, audio, video, and text inputs. The model began rolling out the same day to Gemini app users with AI Plus, Pro, and Ultra subscriptions, YouTube Shorts, and the YouTube Create app.

Model capabilities and limits

Gemini Omni Flash generates video clips capped at 10 seconds at launch. According to Koray Kavukcuoglu, CTO of Google DeepMind, the model supports conversational editing where each instruction builds on previous ones while preserving character identity and scene continuity across multiple turns.

Google claims the model has improved physics simulation including gravity, kinetic energy, and fluid dynamics. The company demonstrated prompts ranging from claymation protein-folding animations to chain-reaction physics tracks, though benchmark scores comparing Omni to Veo 3 or competing models like ByteDance's Seedance have not been disclosed.

The 10-second limit is shorter than OpenAI's Sora, which generates clips up to 60 seconds. Google has not disclosed per-clip costs, compute footprint per generation, or pricing for tiers beyond Flash.

Speech editing withheld, SynthID mandatory

Google is explicitly withholding general-purpose audio and speech editing capabilities. "We are still working to test this and better understand how we can bring this capability to users responsibly," Kavukcuoglu wrote, in what appears to be a deliberate step back from deepfake territory.

The model includes digital avatar generation requiring users to record their voice and likeness by speaking a series of numbers aloud during onboarding.

All videos generated with Omni carry Google's SynthID imperceptible digital watermark by default. Users can verify whether a clip was generated by Omni through the Gemini app, Gemini in Chrome, and Google Search. The SynthID implementation follows the C2PA open standard that OpenAI adopted earlier this year.

Availability and rollout

API access for developers and enterprise customers will launch in the coming weeks. Google has not disclosed the underlying model architecture relative to Veo 3, benchmark evaluation methodology, or timeline for enabling speech editing across the Omni family.

The model launched alongside Gemini 3.5 and other announcements at I/O 2026 that Sundar Pichai framed as the "agentic Gemini era" in his keynote.

What this means

Google's decision to withhold speech editing while releasing video generation marks a conservative deployment strategy compared to frontier competitors. The 10-second clip limit and lack of disclosed pricing create uncertainty about whether Omni represents a technical advance or primarily a product integration of existing capabilities. The mandatory SynthID watermarking positions Google as prioritizing provenance over feature completeness, though the effectiveness of imperceptible watermarks against adversarial removal remains an open question. The API rollout in coming weeks will reveal whether the cost structure and extended clip lengths under paid tiers can compete with Sora's current specifications.

Related Articles

model release

Qwen Launches Qwen3.8 27B, an Open-Weight Vision-Language Model with 262K Context

Qwen has released Qwen3.8 27B, a 27-billion-parameter dense vision-language model with a 262K token context window, available now via OpenRouter at $0.45 per million input tokens and $3.20 per million output tokens.

model release

Liquid AI Releases LFM2.5-VL-3B, a 3B-Class Vision-Language Model Built for On-Device Deployment

Liquid AI has released LFM2.5-VL-3B, a multimodal upgrade to its LFM2-VL-3B model built for on-device grounding, object detection, and document OCR. The model runs at 228 tokens/sec on an Apple M5 Max and 116 tokens/sec on an AMD Ryzen AI Max+ 395, using under 3.3 GB of memory.

model release

Google DeepMind Ships Gemini 3.7 Flash, Closing Gap With Claude 4.8 and GPT-5.5

Google DeepMind has released Gemini 3.7 Flash, a new entry in its fast-tier model line that reportedly closes a performance gap that opened up under Gemini 3.5 and 3.6 Flash against Anthropic's Claude 4.8+ and OpenAI's GPT-5.5+ series. Full pricing and benchmark details have not yet been disclosed.

model release

Google Releases Gemini 3.7 Flash, Cuts Price in Half Versus 3.6 Flash

Google has released Gemini 3.7 Flash, just three weeks after Gemini 3.6 Flash, claiming substantial gains in coding, web development, and document reasoning. The model launches at an introductory price of $0.75 per 1M input tokens and $3.75 per 1M output tokens — half the cost of its predecessor.

Comments

Loading...