Black Forest Labs Launches FLUX 3 Video, Claims It Beats Seedance 2.0 on Elo Rankings
Black Forest Labs has made FLUX 3 Video generally available via its API, offering up to 20-second HD/Full HD clips with native audio and lip-sync in 14+ languages. The company claims its internal Elo benchmarks put the model ahead of Seedance 2.0, Gemini Omni Flash, and Minimax H3.
Black Forest Labs (BFL) has made FLUX 3 Video generally available through its API and select partners, expanding the company's lineup beyond image generation into full video-with-audio synthesis.
The model produces HD and Full HD clips up to 20 seconds long, with native audio generation covering dialogue, sound effects, and ambient noise baked directly into the output. FLUX 3 supports text-to-video, image-to-video, keyframe-based generation, video continuation, and multiple scenes or camera angles within a single clip.
BFL says the model can render typography directly within scenes, follow complex multi-part prompts, and draw on world knowledge for applications like documentary-style content. It also generates lip-synced dialogue in more than 14 languages, a feature aimed at localization and multilingual content production.
Benchmark claims
According to BFL's own internal testing, FLUX 3 tops its Elo rankings with a score of 1,135 for text-to-video and 1,051 for image-to-video. The company claims these scores place FLUX 3 ahead of Google's Gemini Omni Flash, Minimax H3, and ByteDance's Seedance 2.0. These figures come from BFL's proprietary evaluation and have not been independently verified by third-party benchmarking organizations.
Pricing
FLUX 3 Video pricing is structured per second of output and split between a lower-cost draft mode and full-quality generation:
Draft mode (HD only):
- Text-to-video / image-to-video: $0.06 per second
- Video-to-video: $0.12 per second
Full quality (HD):
- Text-to-video / image-to-video: $0.17 per second
- Video-to-video: $0.41 per second
Full quality (Full HD):
- Text-to-video / image-to-video: $0.29 per second
- Video-to-video: $0.53 per second
Audio generation is included in all listed prices at no extra cost. A 10-second Full HD text-to-video clip at full quality would run $2.90, while the same clip in draft mode would cost $0.60.
What this means
BFL is positioning FLUX 3 Video as a direct competitor to Seedance 2.0, Gemini's video tools, and Minimax's offerings in a market where synchronized audio-video generation is becoming the baseline expectation rather than a differentiator. The company's Elo comparisons are self-reported, so independent verification from users or third-party leaderboards will matter more than BFL's internal numbers.
The tiered pricing — draft mode versus full quality, HD versus Full HD — signals that BFL expects most usage to be iterative: cheap drafts for testing prompts, full-quality renders for final output. At $0.41-$0.53 per second for video-to-video at full quality, costs can add up quickly for longer-form content, which may push production studios toward draft mode for early-stage work and reserve full quality for final delivery. The 20-second cap and 14-language lip-sync support suggest BFL is targeting short-form marketing, social content, and localized dialogue-heavy clips rather than long-form film production.
Related Articles
Qwen3.8-Omni-Flash Prices Multimodal AI at $0.15/$0.47 per Million Tokens, Undercutting Gemini Flash by 5x
Alibaba's Qwen team released Qwen3.8-Omni-Flash, a multimodal model for AI agents that processes audio and video with a 1 million token context window. Pricing undercuts Google's Gemini 3.8 Flash by roughly 5x on input and 8x on output, according to Qwen.
PrismML Releases Ternary Bonsai 2 27B, a Compressed Reasoning Model with 262K Context
PrismML has released Ternary Bonsai 2 27B, a 27B-parameter reasoning model derived from Qwen3.8-27B that uses ternary weight compression to shrink to roughly 8.5 GB. The model supports a 262K-token context window, image understanding, tool calling, and thinks by default at 'xhigh' reasoning effort.
OpenAI RLHF Co-Inventor Launches Jev, a Non-LLM Model That Outputs Probabilities Instead of Text
TypeSafe AI, founded by RLHF co-inventor Diogo Almeida, has released Jev, a transformer-based model that outputs probabilities rather than text. Developers report it running 5 to 20 times cheaper and faster than LLMs for classification tasks.
Z.ai Releases GLM-5.3-FlashX, a 200 Tokens/Second Variant of Its GLM-5.3-Flash Model
Z.ai has released GLM-5.3-FlashX, a high-speed variant of GLM-5.3-Flash built on a hybrid sparse and linear attention architecture with 320B total parameters (18B active). The model supports a 1M-token context window and claims inference speeds of up to 200 tokens per second.
Comments
Loading...