model releaseByteDance

ByteDance's Seedance 2.5 Generates 30-Second AI Video Clips With Synced Audio

TL;DR

ByteDance released Seedance 2.5, an AI video model that generates synchronized video and audio in a single pass, producing clips up to 30 seconds long that can be extended further. That's roughly triple the length of Google's Gemini Omni Flash.

2 min read
0

ByteDance has released Seedance 2.5, an updated version of its AI video generation model that produces synchronized video and audio in a single generation pass. The model creates clips up to 30 seconds long, which can be extended multiple times for longer sequences.

By comparison, Google's Gemini Omni Flash tops out at roughly 10 seconds per clip. ByteDance has not disclosed pricing or API rate limits for Seedance 2.5.

What's new

Seedance 2.5 accepts a broader set of reference inputs than its predecessor. Users can upload up to 30 images, 10 video clips, and 10 audio files as references, allowing the model to construct scenes with multiple characters and varying camera angles from those inputs. ByteDance also says textures, lighting, and skin detail have improved over Seedance 2.0, though the company has not published quantitative benchmark comparisons for these visual quality claims.

To demonstrate the model, ByteDance produced a short film titled "The Missing Pair" using Seedance 2.5 exclusively, showcasing multi-character scenes and camera work generated by the system.

Availability

Seedance 2.5 is live now on ByteDance's Jimeng AI and Doubao Pro platforms. API access through BytePlus ModelArk is planned but not yet available, so third-party developers cannot currently integrate the model into external products.

Track record

The previous generation, Seedance 2.0, currently leads the image-to-video leaderboard among models with audio generation, according to independent benchmarking site Artificial Analysis. That model gained attention in the film community after "District 9" director Neill Blomkamp used it to create "Nightborne," a 13-minute short film generated entirely with AI — one of the longer AI-generated narrative works produced to date.

What this means

The 30-second ceiling matters less for raw duration than for what it enables: fewer stitching seams. Combining separately generated clips into a coherent scene is one of the persistent pain points in AI video production, since lighting, character consistency, and audio sync tend to drift between segments. A single-pass 30-second generation with native audio reduces how many of those seams a production team has to manually patch.

The expanded reference inputs — 30 images, 10 videos, 10 audio files — point toward ByteDance targeting professional ad and content studios rather than casual users experimenting with text prompts. That's a different competitive lane than OpenAI's Sora or Runway's consumer-facing tools, and it puts Seedance closer to being a production tool than a novelty generator.

Still, the comparison to Gemini Omni Flash's 10-second limit should be read carefully: longer native generation often trades off against per-second visual fidelity or coherence, and ByteDance has not released independent benchmark data to confirm quality holds steady across the full 30 seconds. Until BytePlus ModelArk API access ships and outside benchmarking firms test the model directly, ByteDance's specific quality claims — improved textures, lighting, skin detail — remain unverified. The Artificial Analysis leaderboard placement for Seedance 2.0 is independently confirmed, which lends some credibility to ByteDance's trajectory, but 2.5 has not yet been through the same third-party evaluation.

Related Articles

model release

InclusionAI Releases Ling 3.0 Flash VL, Adding Vision to Its 124B MoE Model

InclusionAI has released Ling 3.0 Flash VL, a vision-language extension of its 124B total-parameter, 5.5B active Mixture-of-Experts model. The model adds native image and video understanding, supports a 131K token context window, and is priced at $0.06 per 1M input tokens and $0.18 per 1M output tokens via OpenRouter.

model release

Alibaba's Qwen Releases Qwen-Drive-1.0-4B, a Unified VLM for Autonomous Driving Perception and Planning

Alibaba's Qwen team has released Qwen-Drive-1.0-4B, a 4B-parameter vision-language model built on Qwen3.5 that unifies 3D perception, driving question answering, and motion planning in one framework. The model reports strong open-loop, pseudo-closed-loop, and closed-loop driving benchmark results while claiming minimal loss of general vision-language ability.

model release

DeepSeek Releases V4.1-Flash: 552B MoE Model Cuts KV Cache to 890 Bytes Per Token

DeepSeek has released V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model supporting 1M-token context and activating only 8B parameters during prefill. The model uses a new Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 to cut global KV cache to 890 bytes per token, roughly a quarter of its predecessor.

model release

AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks

Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.

Comments

Loading...