video generation
14 articles tagged with video generation
Fal Launches H3 Max Live, a Post-Trained Minimax H3 Variant That Generates Video Faster Than Real Time
Fal released H3 Max Live, a post-trained and inference-optimized version of Minimax's H3 video model that Fal claims runs up to 35x faster than the official endpoint. The model generates video faster than it can be watched, enabling an infinite, chat-directed live video stream.
Alibaba Launches Wan3.0, Generating AI Videos Up to 30 Seconds From Text, Images, and Documents
Alibaba's Wan3.0 video generation model is now in beta, producing clips up to 30 seconds long and accepting text, images, video, audio, and documents like PDFs and PowerPoint files as input. Pricing runs per second of output across two tiers and three resolutions.
Vercel AI SDK Adds Support for xAI's Grok Imagine Video 1.5, Including 1080p Output
The @ai-sdk/xai package version 3.0.121 adds support for xAI's grok-imagine-video-1.5 model, including native 1080p resolution for text-to-video and image-to-video generation. The update also introduces referenceVoiceIds for reference audio and fixes a routing bug in reference-to-video mode.
Vercel AI SDK Adds Support for xAI's Grok Imagine Video 1.5, Including 1080p Generation
Vercel released @ai-sdk/xai version 4.0.36, adding support for xAI's grok-imagine-video-1.5 model with native 1080p resolution and a new referenceVoiceIds parameter for reference-to-video audio. The update also fixes a routing bug that misdirected certain video generation requests.
Lightricks Releases LTX-2.5, a 22B-Parameter Open-Weight Video and Audio World Model
Lightricks has released LTX-2.5, an open-weight world model that generates synchronized video and audio from text, image, and video inputs. The 22B-parameter model adds native multishot generation, a new diffusion video decoder, and a custom Gemma4 12B text encoder.
Black Forest Labs Launches FLUX 3 Video, Claims It Beats Seedance 2.0 on Elo Rankings
Black Forest Labs has made FLUX 3 Video generally available via its API, offering up to 20-second HD/Full HD clips with native audio and lip-sync in 14+ languages. The company claims its internal Elo benchmarks put the model ahead of Seedance 2.0, Gemini Omni Flash, and Minimax H3.
Black Forest Labs Announces FLUX 3 Video, First Details Cover Generation Capabilities
Black Forest Labs has published the first part of its FLUX 3 Video release notes, focused on the model's generation capabilities. Full technical specifications, pricing, and benchmark data have not yet been disclosed.
MiniMax H3 Becomes First Open Video Model to Top an AI Video Ranking
MiniMax has released open weights for H3, a 33-billion-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis — the first time an open model has topped a video generation category. The model accepts text, images, video, and audio in a single prompt, though its highest-resolution module remains closed.
MiniMax Releases H3, a 33B-Parameter Omni-Modal Model That Generates 2K Video With Native Stereo Audio
MiniMax has published MiniMax-H3, a 33-billion-parameter omni-modal generative model capable of producing up to 15 seconds of 2K video with native stereo audio. The model accepts text, image, video, and audio inputs, though its full 2K pipeline depends on a hosted preprocessing component not included in the open-source release.
Black Forest Labs Releases Flux 3, Its First Model to Generate Video With Native Audio Up to 20 Seconds
Black Forest Labs has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio that generates videos up to 20 seconds long with synchronized native audio. The company also introduced Flux-mimic, a robotics action model already being tested at Audi.
Vercel AI SDK adds Grok 4.5 model support and video reference inputs
Vercel released AI SDK version 3.0.106 for xAI integration, adding support for the Grok 4.5 model identifier and expanding input reference capabilities to include video files for reference-to-video generation workflows.
Google releases Nano Banana 2 Lite: 4-second image generation at $0.034 per 1,000 images
Google released Nano Banana 2 Lite, an AI image generator that produces images in 4 seconds and costs $0.034 per 1,000 images. The model is optimized for high-volume workflows and replaces the original Nano Banana as Google's entry-level image generation offering.
Google launches Gemini Omni Flash, multimodal video generation model available to AI Plus subscribers
Google has released Gemini Omni Flash, the first model in its new Gemini Omni family designed to generate video content from text, images, video, and audio inputs. The model is available now to AI Plus subscribers, with free access coming to YouTube Shorts and YouTube Create later this week.
LPM 1.0 generates 45-minute real-time lip-synced video from single photo, no public release planned
Researchers have introduced LPM 1.0, an AI model that generates real-time video of a speaking, listening, or singing character from a single image, with lip-synced speech and facial expressions stable for up to 45 minutes. The system integrates directly with voice AI models like ChatGPT but remains a research project with no planned public release.