ByteDance's Seedance 2.5 Generates 30-Second AI Video Clips With Synced Audio
ByteDance released Seedance 2.5, an AI video model that generates synchronized video and audio in a single pass, producing clips up to 30 seconds long that can be extended further. That's roughly triple the length of Google's Gemini Omni Flash.
ByteDance has released Seedance 2.5, an updated version of its AI video generation model that produces synchronized video and audio in a single generation pass. The model creates clips up to 30 seconds long, which can be extended multiple times for longer sequences.
By comparison, Google's Gemini Omni Flash tops out at roughly 10 seconds per clip. ByteDance has not disclosed pricing or API rate limits for Seedance 2.5.
What's new
Seedance 2.5 accepts a broader set of reference inputs than its predecessor. Users can upload up to 30 images, 10 video clips, and 10 audio files as references, allowing the model to construct scenes with multiple characters and varying camera angles from those inputs. ByteDance also says textures, lighting, and skin detail have improved over Seedance 2.0, though the company has not published quantitative benchmark comparisons for these visual quality claims.
To demonstrate the model, ByteDance produced a short film titled "The Missing Pair" using Seedance 2.5 exclusively, showcasing multi-character scenes and camera work generated by the system.
Availability
Seedance 2.5 is live now on ByteDance's Jimeng AI and Doubao Pro platforms. API access through BytePlus ModelArk is planned but not yet available, so third-party developers cannot currently integrate the model into external products.
Track record
The previous generation, Seedance 2.0, currently leads the image-to-video leaderboard among models with audio generation, according to independent benchmarking site Artificial Analysis. That model gained attention in the film community after "District 9" director Neill Blomkamp used it to create "Nightborne," a 13-minute short film generated entirely with AI — one of the longer AI-generated narrative works produced to date.
What this means
The 30-second ceiling matters less for raw duration than for what it enables: fewer stitching seams. Combining separately generated clips into a coherent scene is one of the persistent pain points in AI video production, since lighting, character consistency, and audio sync tend to drift between segments. A single-pass 30-second generation with native audio reduces how many of those seams a production team has to manually patch.
The expanded reference inputs — 30 images, 10 videos, 10 audio files — point toward ByteDance targeting professional ad and content studios rather than casual users experimenting with text prompts. That's a different competitive lane than OpenAI's Sora or Runway's consumer-facing tools, and it puts Seedance closer to being a production tool than a novelty generator.
Still, the comparison to Gemini Omni Flash's 10-second limit should be read carefully: longer native generation often trades off against per-second visual fidelity or coherence, and ByteDance has not released independent benchmark data to confirm quality holds steady across the full 30 seconds. Until BytePlus ModelArk API access ships and outside benchmarking firms test the model directly, ByteDance's specific quality claims — improved textures, lighting, skin detail — remain unverified. The Artificial Analysis leaderboard placement for Seedance 2.0 is independently confirmed, which lends some credibility to ByteDance's trajectory, but 2.5 has not yet been through the same third-party evaluation.
Related Articles
InclusionAI Releases Ling 3.0 Flash VL, Adding Vision to Its 124B MoE Model
InclusionAI has released Ling 3.0 Flash VL, a vision-language extension of its 124B total-parameter, 5.5B active Mixture-of-Experts model. The model adds native image and video understanding, supports a 131K token context window, and is priced at $0.06 per 1M input tokens and $0.18 per 1M output tokens via OpenRouter.
Alibaba's Qwen Releases Qwen-Drive-1.0-4B, a Unified VLM for Autonomous Driving Perception and Planning
Alibaba's Qwen team has released Qwen-Drive-1.0-4B, a 4B-parameter vision-language model built on Qwen3.5 that unifies 3D perception, driving question answering, and motion planning in one framework. The model reports strong open-loop, pseudo-closed-loop, and closed-loop driving benchmark results while claiming minimal loss of general vision-language ability.
DeepSeek Releases V4.1-Flash: 552B MoE Model Cuts KV Cache to 890 Bytes Per Token
DeepSeek has released V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model supporting 1M-token context and activating only 8B parameters during prefill. The model uses a new Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 to cut global KV cache to 890 bytes per token, roughly a quarter of its predecessor.
AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks
Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.
Comments
Loading...