model releaseBlack Forest Labs

Black Forest Labs Releases Flux 3, Its First Model to Generate Video With Native Audio Up to 20 Seconds

TL;DR

Black Forest Labs has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio that generates videos up to 20 seconds long with synchronized native audio. The company also introduced Flux-mimic, a robotics action model already being tested at Audi.

3 min read
0

German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio. The headline feature: Flux 3 can generate videos up to 20 seconds long with native, synchronized audio — a first for the company.

What Flux 3 does

Flux 3 supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven chaining of clips into longer multi-shot sequences. BFL claims the model is particularly strong at rendering human facial expressions and matching audio to physical events on screen.

The company describes Flux 3 as a step toward what it calls "real-world visual intelligence" — models that can perceive, predict, and act across physical and digital environments. BFL's underlying argument: no single modality captures reality fully. Images encode spatial structure, video captures change over time, and audio reveals connections between mechanical events and the sounds they produce. Training on all three simultaneously, according to BFL, lets the model fill gaps that single-modality training leaves open.

Benchmark claims against rivals

In early internal evaluations using 10-second clips at 720p resolution, BFL reports the following preference rates for Flux 3 over competing video models:

  • Luma Ray 3.2: 93%
  • Runway Gen-4.5: 77%
  • Grok Imagine Video: 69%
  • Kling v3 Pro: 60%
  • Happy Horse v1: 59%
  • Happy Horse 1.1: 57%
  • Seedance 2.0: 52%
  • Gemini Omni Flash: 52%

A score of 50% indicates a tie. These figures come entirely from BFL's own testing — no independent verification is currently available. The narrowing margins against Seedance 2.0 and Gemini Omni Flash suggest Flux 3 is roughly on par with, rather than clearly ahead of, the strongest existing systems. BFL itself acknowledges the results are preliminary.

Architecture: Self-Flow

Flux 3 is built on BFL's Self-Flow approach, which trains a single model to both generate and understand content across modalities at once. A multimodal transformer uses dedicated encoder and decoder components for images, video, audio, and — notably — actions, converting each into a shared internal representation before producing outputs. BFL claims this unified training process outperforms the flow-matching method the company previously relied on, both in output quality and in the model's understanding of physical dynamics.

Robotics: Flux-mimic

The action component underpins Flux-mimic, a video-action model BFL developed with Mimic Robotics. According to BFL, Flux-mimic is already being tested on production tasks at Audi, extending the model's world-modeling capabilities into robotic control rather than just content generation.

Rollout plan

BFL is staging the release. Flux 3 Video is available now; Flux 3 Image is expected within the next few weeks, with BFL claiming improvements in complex prompt handling and multilingual text rendering. Action prediction will initially roll out through select partners only. BFL also plans an open-weight release of the multimodal backbone, called Flux 3 Dev, though no date has been given. Pricing has not been disclosed for any tier.

What this means

Flux 3 pushes BFL beyond static image generation into video-audio-action territory occupied by Google, Runway, Luma, and ByteDance's Seedance. The self-reported benchmarks — while unverified — indicate BFL is at least competitive with, rather than trailing, the current leaders in video generation. The more consequential piece may be Flux-mimic: pairing a world model with robotics testing at a major manufacturer like Audi signals BFL is betting on physical-world applications, not just content creation, as its next growth area. Until independent benchmarks and pricing appear, treat the comparative claims as directional rather than definitive.

Related Articles

model release

Google Releases Gemini Omni Flash Preview, a Multimodal Model for 720p Video Generation

Google has released Gemini Omni Flash Preview, a native multimodal model that generates short 720p videos with native audio from text, image, and video inputs. The model is available now via OpenRouter with a 131K token context window.

model release

Black Forest Labs Unveils FLUX.2 [klein]: A Distilled Model for Interactive Image Generation

Black Forest Labs has released FLUX.2 [klein], a lightweight variant of its FLUX.2 image generation model family designed for faster, more interactive use. The company frames the release as a step toward 'interactive visual intelligence,' though detailed benchmarks and pricing have not yet been disclosed.

model release

Alibaba Releases Qwen-Image-3.0, an Image Generator That Renders 10-Pixel Text and 3x3 Infographic Grids in One Pass

Alibaba's Qwen team has released Qwen-Image-3.0, an image generator that accepts prompts up to 4,500 tokens and can render legible text as small as ten pixels, complex LaTeX formulas, and twelve languages in a single pass. The model is currently invite-only via API, and unlike its predecessor, it likely won't ship with open weights.

model release

Anthropic Launches Claude Opus 5, Claims Parity With Rival Fable 5 at Half the Cost

Anthropic has released Claude Opus 5, its new flagship model, claiming performance comparable to rival model Fable 5 at half the cost. The company says Opus 5 leads on several coding and knowledge-work benchmarks while requiring far less manual intervention.

Comments

Loading...