Black Forest Labs Releases FLUX 3 Action, a 7B Open-Weights World Action Model, Claims Top RoboLab Benchmark Score
Black Forest Labs has released FLUX 3 Action, a 7B parameter open-weights World Action Model. The company claims it achieves first place on the RoboLab benchmark, though independent verification is pending.
Black Forest Labs has released FLUX 3 Action, a 7 billion parameter open-weights World Action Model. According to Black Forest Labs, the model achieves first place on the RoboLab benchmark, a evaluation suite used to measure performance on embodied and physical-action tasks.
FLUX 3 Action extends the FLUX line beyond image generation into the domain of "world action models" — systems designed to predict and generate action sequences grounded in a model of the physical or simulated world. This positions the release as a departure from prior FLUX models, which have focused primarily on text-to-image generation, and places Black Forest Labs in more direct competition with robotics and embodied-AI labs building similar world models.
What's known
- Parameter count: 7 billion
- Weights: Open weights, available for download and self-hosting
- Benchmark claim: First place on RoboLab, according to Black Forest Labs
Black Forest Labs has not yet disclosed the model's context window, training data cutoff, licensing terms for commercial use, or hosted API pricing. The company also has not published a technical report detailing the RoboLab scores relative to competing models, so the "first place" claim currently rests on Black Forest Labs' own announcement rather than independent third-party evaluation.
Why a 7B open model matters here
At 7B parameters, FLUX 3 Action is small enough to run on a single high-end GPU, which matters for robotics and embodied-AI applications where inference often needs to happen on-device or with low latency rather than through a cloud API. Open weights also let researchers and robotics companies fine-tune the model on proprietary action data — simulator logs, teleoperation datasets, or real-world trajectories — without depending on a hosted service.
The RoboLab benchmark itself is less established in the broader AI benchmark landscape than suites like MMLU or HumanEval, so the significance of a first-place finish will depend on which models and labs are included in the comparison set. Until Black Forest Labs or a third party publishes a full leaderboard with methodology, the claim should be treated as a company-reported result.
What this means
FLUX 3 Action signals that Black Forest Labs is expanding from generative imagery into the world-model and robotics space, where labs like Google DeepMind, NVIDIA, and various robotics startups are already competing to build models that translate perception into physical action. A 7B open-weights model is a low-cost entry point for developers who want to experiment with embodied AI without training from scratch. But with no independent benchmark verification, no pricing details, and no technical report yet available, the practical utility of FLUX 3 Action for production robotics systems remains unproven outside of Black Forest Labs' own claims. Expect scrutiny from the robotics research community once the weights are downloadable and third parties can run their own RoboLab and out-of-distribution tests.
Related Articles
Black Forest Labs Releases FLUX 3 Action, a 7B-Parameter Open Robotics Model
Black Forest Labs has released FLUX 3 Action, an open-weight robotics model built on its FLUX 3 multimodal foundation. The 7-billion-parameter model reads multi-camera video feeds and predicts what a robot should do next, claiming a record success rate on the RoboLab-120 leaderboard while running nearly 4x faster than the previous best open model.
Google DeepMind's New Chief Prioritizes Fast Gemini 4 Release Over AGI Debate
Google DeepMind's new head Koray Kavukcuoglu says Gemini 4 is in early post-training and could ship well before year-end, following the quiet cancellation of Gemini 3.5 Pro. He downplayed the AGI question that defined predecessor Demis Hassabis's tenure, calling it 'not the right conversation.'
NVIDIA Releases Nemotron 3 Diarization, an Open-Weight Speaker ID Model Supporting Up to 8 Speakers
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model that determines "who spoke when" in audio, supporting both streaming and offline inference for up to eight speakers. The model achieves input buffer latency as low as 80 milliseconds and is available for commercial and non-commercial use.
Fireworks Releases Ember-1, a Reasoning Model That Cuts Token Usage 40% Versus Its Kimi K3 Base
Fireworks Research has released Ember-1, a reasoning model built on Kimi K3 that produces shorter reasoning traces while claiming comparable output quality. The model offers a 1 million token context window at $3 per 1M input tokens and $15 per 1M output tokens.
Comments
Loading...