model releaseBlack Forest Labs

Black Forest Labs Launches FLUX 3 Video, Claims It Beats Seedance 2.0 on Elo Rankings

TL;DR

Black Forest Labs has made FLUX 3 Video generally available via its API, offering up to 20-second HD/Full HD clips with native audio and lip-sync in 14+ languages. The company claims its internal Elo benchmarks put the model ahead of Seedance 2.0, Gemini Omni Flash, and Minimax H3.

2 min read
0

Black Forest Labs (BFL) has made FLUX 3 Video generally available through its API and select partners, expanding the company's lineup beyond image generation into full video-with-audio synthesis.

The model produces HD and Full HD clips up to 20 seconds long, with native audio generation covering dialogue, sound effects, and ambient noise baked directly into the output. FLUX 3 supports text-to-video, image-to-video, keyframe-based generation, video continuation, and multiple scenes or camera angles within a single clip.

BFL says the model can render typography directly within scenes, follow complex multi-part prompts, and draw on world knowledge for applications like documentary-style content. It also generates lip-synced dialogue in more than 14 languages, a feature aimed at localization and multilingual content production.

Benchmark claims

According to BFL's own internal testing, FLUX 3 tops its Elo rankings with a score of 1,135 for text-to-video and 1,051 for image-to-video. The company claims these scores place FLUX 3 ahead of Google's Gemini Omni Flash, Minimax H3, and ByteDance's Seedance 2.0. These figures come from BFL's proprietary evaluation and have not been independently verified by third-party benchmarking organizations.

Pricing

FLUX 3 Video pricing is structured per second of output and split between a lower-cost draft mode and full-quality generation:

Draft mode (HD only):

  • Text-to-video / image-to-video: $0.06 per second
  • Video-to-video: $0.12 per second

Full quality (HD):

  • Text-to-video / image-to-video: $0.17 per second
  • Video-to-video: $0.41 per second

Full quality (Full HD):

  • Text-to-video / image-to-video: $0.29 per second
  • Video-to-video: $0.53 per second

Audio generation is included in all listed prices at no extra cost. A 10-second Full HD text-to-video clip at full quality would run $2.90, while the same clip in draft mode would cost $0.60.

What this means

BFL is positioning FLUX 3 Video as a direct competitor to Seedance 2.0, Gemini's video tools, and Minimax's offerings in a market where synchronized audio-video generation is becoming the baseline expectation rather than a differentiator. The company's Elo comparisons are self-reported, so independent verification from users or third-party leaderboards will matter more than BFL's internal numbers.

The tiered pricing — draft mode versus full quality, HD versus Full HD — signals that BFL expects most usage to be iterative: cheap drafts for testing prompts, full-quality renders for final output. At $0.41-$0.53 per second for video-to-video at full quality, costs can add up quickly for longer-form content, which may push production studios toward draft mode for early-stage work and reserve full quality for final delivery. The 20-second cap and 14-language lip-sync support suggest BFL is targeting short-form marketing, social content, and localized dialogue-heavy clips rather than long-form film production.

Related Articles

model release

Black Forest Labs Announces FLUX 3 Video, First Details Cover Generation Capabilities

Black Forest Labs has published the first part of its FLUX 3 Video release notes, focused on the model's generation capabilities. Full technical specifications, pricing, and benchmark data have not yet been disclosed.

model release

MiniMax H3 Becomes First Open Video Model to Top an AI Video Ranking

MiniMax has released open weights for H3, a 33-billion-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis — the first time an open model has topped a video generation category. The model accepts text, images, video, and audio in a single prompt, though its highest-resolution module remains closed.

model release

ByteDance's Seedance 2.5 Generates 30-Second AI Video Clips With Synced Audio

ByteDance released Seedance 2.5, an AI video model that generates synchronized video and audio in a single pass, producing clips up to 30 seconds long that can be extended further. That's roughly triple the length of Google's Gemini Omni Flash.

model release

MiniMax Releases H3, a 33B-Parameter Omni-Modal Model That Generates 2K Video With Native Stereo Audio

MiniMax has published MiniMax-H3, a 33-billion-parameter omni-modal generative model capable of producing up to 15 seconds of 2K video with native stereo audio. The model accepts text, image, video, and audio inputs, though its full 2K pipeline depends on a hosted preprocessing component not included in the open-source release.

Comments

Loading...