model releaseBlack Forest Labs

Black Forest Labs Releases Flux 3, Its First Model to Generate Video With Native Audio Up to 20 Seconds

TL;DR

Black Forest Labs has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio that generates videos up to 20 seconds long with synchronized native audio. The company also introduced Flux-mimic, a robotics action model already being tested at Audi.

3 min read
0

German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model trained jointly on images, video, and audio. The headline feature: Flux 3 can generate videos up to 20 seconds long with native, synchronized audio — a first for the company.

What Flux 3 does

Flux 3 supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven chaining of clips into longer multi-shot sequences. BFL claims the model is particularly strong at rendering human facial expressions and matching audio to physical events on screen.

The company describes Flux 3 as a step toward what it calls "real-world visual intelligence" — models that can perceive, predict, and act across physical and digital environments. BFL's underlying argument: no single modality captures reality fully. Images encode spatial structure, video captures change over time, and audio reveals connections between mechanical events and the sounds they produce. Training on all three simultaneously, according to BFL, lets the model fill gaps that single-modality training leaves open.

Benchmark claims against rivals

In early internal evaluations using 10-second clips at 720p resolution, BFL reports the following preference rates for Flux 3 over competing video models:

  • Luma Ray 3.2: 93%
  • Runway Gen-4.5: 77%
  • Grok Imagine Video: 69%
  • Kling v3 Pro: 60%
  • Happy Horse v1: 59%
  • Happy Horse 1.1: 57%
  • Seedance 2.0: 52%
  • Gemini Omni Flash: 52%

A score of 50% indicates a tie. These figures come entirely from BFL's own testing — no independent verification is currently available. The narrowing margins against Seedance 2.0 and Gemini Omni Flash suggest Flux 3 is roughly on par with, rather than clearly ahead of, the strongest existing systems. BFL itself acknowledges the results are preliminary.

Architecture: Self-Flow

Flux 3 is built on BFL's Self-Flow approach, which trains a single model to both generate and understand content across modalities at once. A multimodal transformer uses dedicated encoder and decoder components for images, video, audio, and — notably — actions, converting each into a shared internal representation before producing outputs. BFL claims this unified training process outperforms the flow-matching method the company previously relied on, both in output quality and in the model's understanding of physical dynamics.

Robotics: Flux-mimic

The action component underpins Flux-mimic, a video-action model BFL developed with Mimic Robotics. According to BFL, Flux-mimic is already being tested on production tasks at Audi, extending the model's world-modeling capabilities into robotic control rather than just content generation.

Rollout plan

BFL is staging the release. Flux 3 Video is available now; Flux 3 Image is expected within the next few weeks, with BFL claiming improvements in complex prompt handling and multilingual text rendering. Action prediction will initially roll out through select partners only. BFL also plans an open-weight release of the multimodal backbone, called Flux 3 Dev, though no date has been given. Pricing has not been disclosed for any tier.

What this means

Flux 3 pushes BFL beyond static image generation into video-audio-action territory occupied by Google, Runway, Luma, and ByteDance's Seedance. The self-reported benchmarks — while unverified — indicate BFL is at least competitive with, rather than trailing, the current leaders in video generation. The more consequential piece may be Flux-mimic: pairing a world model with robotics testing at a major manufacturer like Audi signals BFL is betting on physical-world applications, not just content creation, as its next growth area. Until independent benchmarks and pricing appear, treat the comparative claims as directional rather than definitive.

Related Articles

model release

Alibaba Releases Qwen-Drive 1.0, an Open Driving Model That Explains Its Own Decisions

Alibaba has released Qwen-Drive 1.0, a driving model built on Qwen3.5-4B that handles spatial perception, route planning, and cockpit dialogue in a single system. Reinforcement learning cut the rate of off-road driving errors in simulation from 24 percent to 12 percent, though the model's stated reasoning doesn't always match its actual maneuvers.

model release

Microsoft Releases VibeVoice-ASR-Streaming-7B, an Open-Weight Streaming Speech Recognition Model with Speaker Attributio

Microsoft Research has released VibeVoice-ASR-Streaming-7B, an open-weight streaming automatic speech recognition model that transcribes both who is speaking and what they say in real time. The model, listed at 9B parameters despite its name, supports 10 languages and custom hotwords under an MIT license.

model release

Google's WeatherNext 3 Drops Physics Simulations, Learns Weather Forecasting Directly From Satellite Data

Google and DeepMind released WeatherNext 3, an AI weather model that trains directly on live geostationary satellite data instead of physics-based simulations. The model produces hourly forecasts at up to 5-kilometer resolution and now powers weather features in Google Search, Maps, and Gemini.

model release

Google Launches Lyria 3.5 AI Music Model Directly Inside the Gemini App

Google has released Lyria 3.5, a new AI music generation model, directly inside the Gemini app alongside availability in AI Studio, Flow Music, and Vids. Google claims the model was trained exclusively on licensed content and produces more expressive vocals than its predecessor.

Comments

Loading...