Google launches Gemini Omni, multimodal AI video generator with avatar cloning and physics modeling
Google has released Gemini Omni, a multimodal AI video generation tool that accepts text, images, audio, and video as inputs. The first tier, Gemini Omni Flash, includes avatar cloning that creates digital versions of users and incorporates physics modeling for realistic motion.
Gemini Omni Flash — Quick Specs
Google launches Gemini Omni, multimodal AI video generator with avatar cloning and physics modeling
Google has released Gemini Omni, a multimodal AI video generation tool that the company positions as doing "for video what Nano Banana did for images." The first tier, Gemini Omni Flash, is now rolling out to the Gemini app, Google Flow, and YouTube Shorts.
Core capabilities
According to Google, Omni accepts four input types: text, images, audio (currently voice recordings only), and video. The company claims the model can "create anything from any input," with plans to expand beyond video generation. The tool incorporates SynthID digital fingerprinting technology to identify AI-generated content.
Avatar cloning feature
Omni includes an Avatars feature that creates digital replicas of users, generating videos that "look and sound like you," according to Google. The company stated it is "still working to test" the capability to edit videos to change audio and speech, citing responsible deployment concerns.
Physics modeling
The model incorporates what Google describes as "an improved intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics." This physics modeling aims to create more realistic motion compared to earlier AI video tools that treated objects like ragdolls rather than physical entities.
Natural language editing
Omni supports conversational video editing through natural language instructions. Google claims that "every instruction builds on the last" while maintaining character consistency and scene continuity. The company said users can "change specific things, or change everything" in existing videos, including adding characters, transforming objects, or altering backgrounds.
Google has not disclosed video resolution limits, supported aspect ratios, maximum clip length, or pricing per plan tier. The company also has not specified whether Omni will integrate with professional editing software like Final Cut, Premiere Pro, or DaVinci Resolve.
Availability
Gemini Omni Flash is rolling out now to enterprise customers through the Gemini app, Google Flow, and YouTube Shorts. Google has not announced whether the web version of Gemini will support Omni or if users must access it through the Flow interface.
What this means
Gemini Omni represents Google's entry into the competitive AI video generation space, directly challenging OpenAI's Sora. The physics modeling and multimodal input capabilities address key weaknesses in earlier AI video tools. However, the lack of disclosed specifications—particularly around resolution, format support, and professional workflow integration—leaves open questions about whether this targets casual creators or professional production environments. The avatar cloning feature introduces significant trust and verification challenges for video content, even with SynthID watermarking.
Related Articles
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Mistral Releases Shieldstral, a 3B Open-Weights Safety Classifier That Matches Models 7x Its Size
Mistral has released Shieldstral, a 3B open-weights safety classifier that reframes content moderation as a policy-adaptive question-answering task. The model claims to match or outperform guard models up to 7x its size on text safety and multimodal benchmarks, and runs on a single 16GB GPU.
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks
Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.
Comments
Loading...