Google Releases Gemini Omni Flash Preview, a Multimodal Model for 720p Video Generation
Google has released Gemini Omni Flash Preview, a native multimodal model that generates short 720p videos with native audio from text, image, and video inputs. The model is available now via OpenRouter with a 131K token context window.
Gemini Omni Flash Preview — Quick Specs
Google has released Gemini Omni Flash Preview, a native multimodal model built for video generation and editing. The model accepts text, images, and video as input and outputs both text and generated video, according to Google.
What the Model Does
Gemini Omni Flash Preview generates short videos at 720p resolution with native audio, supporting text-to-video generation according to Google's model description on OpenRouter. The model's modality is listed as text+image+video → text+video, positioning it as a unified system for both understanding and generating video content rather than a separate pipeline for each task.
The model ships with a 131,072 token context window, giving it room to process substantial text, image, and video inputs in a single request.
Pricing and Availability
Gemini Omni Flash Preview is available now through OpenRouter's API as google/gemini-omni-flash-preview. Pricing is set at $1.50 per million input tokens and $9.00 per million output tokens — a notably higher output cost than Google's text-only Flash models, reflecting the compute demands of video generation.
As a "preview" release, the model carries Google's standard caveat that specifications, pricing, and behavior may change before a stable release. Google has not disclosed a training data cutoff date for this model, nor released benchmark scores against standard video-generation evaluation suites such as VBench or EvalCrafter.
What This Means
Gemini Omni Flash Preview signals Google's push to fold video generation into its mainline Gemini family rather than keeping it siloed in a separate product like Veo. Bundling text, image, and video understanding with video output in one model — accessible through the same API surface as other Gemini models — lowers the integration burden for developers who want to add video generation to existing Gemini-based applications.
The $9/M output pricing is steep relative to text generation but likely reflects the cost of producing video frames plus synchronized audio, a combination that few competitors offer in a single API call. OpenAI's Sora and other dedicated video models typically charge per-second or per-clip rather than per-token, making direct pricing comparisons difficult until more usage data becomes available.
Because this is a preview model with no independent benchmark scores yet, developers should treat output quality and reliability as unverified until third-party testing emerges. The lack of a disclosed training cutoff also means users can't yet assess how current the model's world knowledge is when generating video content tied to recent events or trends.
Related Articles
Black Forest Labs Releases Flux 3, Its First Model to Generate Video With Native Audio Up to 20 Seconds
Black Forest Labs has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio that generates videos up to 20 seconds long with synchronized native audio. The company also introduced Flux-mimic, a robotics action model already being tested at Audi.
Anthropic Launches Claude Opus 5 (Fast) at $10/$50 per Million Tokens, 1M Context Window
Anthropic has released Claude Opus 5 (Fast), a higher-throughput variant of Opus 5 that carries identical capabilities but runs at roughly 2x the price of the standard model. The model ships with a 1 million token context window and is available now through OpenRouter.
Alibaba Releases Qwen-Image-3.0, an Image Generator That Renders 10-Pixel Text and 3x3 Infographic Grids in One Pass
Alibaba's Qwen team has released Qwen-Image-3.0, an image generator that accepts prompts up to 4,500 tokens and can render legible text as small as ten pixels, complex LaTeX formulas, and twelve languages in a single pass. The model is currently invite-only via API, and unlike its predecessor, it likely won't ship with open weights.
Google Launches Gemini Nano 4 and Gemini Intelligence on Samsung's Galaxy Z Fold 8, Flip 8
Samsung's Galaxy Z Fold 8, Fold 8 Ultra, and Flip 8 are the first devices to ship with Google's Gemini Nano 4 on-device model and the new Gemini Intelligence feature tier. The launch comes with strict hardware requirements including 12GB+ RAM and qualified system-on-chips.
Comments
Loading...