Video Models Steer Robots: Black Forest Labs’ Flux 3 Action applies video expertise to a World-Action Model

Many companies that build image and video generators have pivoted to robotics, extending their expertise in world models to action.

Share
Colored blocks on a grid under a robotic arm, labeled left and right, reflect robotics advancements.

Many companies that build image and video generators have pivoted to robotics, extending their expertise in world models to action. A recent model tops Nvidia’s RoboLab simulation benchmark, but the race is close, and none of that benchmark’s leaders have appeared yet on RoboArena, which tests models on real robots.

What’s new: Black Forest Labs (BFL), best known for its FLUX image generators, released FLUX 3 Action, a 7 billion-parameter model that turns camera feeds and text instructions into robot movements. It ranks first on RoboLab-120, completing 42.9 percent of 1,200 trials. The base model and versions fine-tuned for two different robot arms (Franka and SO-101) are available on Hugging Face under BFL’s FLUX Kommunity License, which limits commercial use to companies earning less than $5 million in annual revenue.

How it works: FLUX 3 Action is a world-action model (WAM). Given camera images and a written instruction, it predicts the robot’s next moves along with what its cameras will see as a result. It also takes in the robot’s current joint positions. That differs from vision-language-action models (VLAs), such as Physical Intelligence’s π0.5 and Nvidia’s GR00T, which predict actions only.

  • Training: BFL built the model on FLUX 3, whose pretraining tokens were more than 95 percent video. The remainder of pretraining tokens came mostly from images, with audio accounting for less than 0.5 percent. In a second training stage focused on actions, 15.93 percent of the samples came from teleoperated robots. The rest came from sources such as video game play and first-person footage of human hands, according to BFL’s research notes.
  • Architecture: The model is a 7 billion parameter diffusion transformer. A frozen video encoder processes camera frames, and a frozen copy of Qwen3-VL-4B processes instructions. Actions form a second stream of tokens that the model denoises together with future video frames.
  • Output: Each call typically returns 32 actions that cover about two seconds of motion, plus optional predicted video. The robot carries out the first few actions, then the model looks again and replans.
  • Adaptation: Each new robot needs its own input and output layers. BFL adapted the model to a low-cost SO-101 arm using about 200 demonstrations. It also fine-tuned the model to play two simple video games and fly a simulated drone. BFL says it sees early promise in simulated vehicle control and computer use too.
  • Results: World-action models, either alone or paired with vision-language models, hold four of the top five spots on RoboLab-120. Well-known VLAs rank lower. π0.5 is ninth at 28.0 percent, and NVIDIA’s GR00T N1.6 is 13th at 7.2 percent.
  • A close race: FLUX 3 Action (42.9 percent) leads HiDream-O1-Embodied (39.9 percent), Atomic-WAM (39.6 percent) and OASIS WAM (39.0 percent). On the benchmark’s hardest tasks, HiDream (32.9 percent) and OASIS (32.4 percent) outperform FLUX 3 Action (28.2 percent).
  • Size: BFL says FLUX 3 Action has less than half the parameters of the previous best open model, Nvidia’s 16 billion-parameter Cosmos3-Nano-Policy, which scored 36.8 percent. However, the leaderboard lists FLUX 3 Action as requiring 69 gigabytes of GPU memory, compared with 40 gigabytes for Cosmos3-Nano-Policy.
  • Speed: BFL says FLUX 3 Action runs 1.52 to 3.95 times faster than Cosmos 3 Nano, depending on the GPU and model version. Its fastest version, which scores 38.3 percent on RoboLab, runs 1.34 to 2.28 times faster than π0.5 because it plans 2.13 seconds of motion per call, compared with π0.5’s 1.0 second. Per call, it’s no faster. On an Nvidia B200 GPU, BFL’s median timings were 32.29 milliseconds for FLUX 3 Action, 31.99 for π0.5, and 320.40 for Cosmos 3 Nano, each in its fastest tested setting.
  • Real robots: In blind tests that BFL arranged with Positronic Robotics, a Franka arm running FLUX 3 Action completed 28 of 30 attempts at 10 single-object DROID tasks. Cosmos 3 Nano completed 27, DreamZero 20, and π0.5 13.

Behind the news: Testing robots in the real world is slow, and every lab’s hardware differs, so the field leans on simulated benchmarks. Some older benchmarks have become too saturated by robots’ success to separate the best models. 

  • Nvidia researchers presented RoboLab in July. Its 120 tabletop tasks, such as stacking blocks in a given order or putting all the green fruit on a plate, are each phrased three ways, from vague to specific. It’s designed for models trained on real-world robot data, so they can’t memorize the test environment, and the leaderboard flags entries that were trained on simulated data. New tasks can be generated in minutes to keep the benchmark from going stale.
  • RoboLab’s authors say its rankings correlate strongly with those of RoboArena, which ranks models by blind head-to-head comparisons on real robot arms at academic labs. However, their paper bases that correlation on just four models. RoboLab’s top five models haven’t yet appeared on RoboArena. A world-action model, DreamZero, ranks second on RoboArena, even though it only places 10th on RoboLab.
  • Nvidia runs RoboLab and tests its own Cosmos and GR00T models. It also worked with BFL on FLUX 3 Action’s fine-tuning recipes and deployment to Nvidia’s Jetson edge computers. BFL credits Nvidia’s Isaac Lab and RoboLab teams with providing the basis for most of its evaluations, and it trained the model on Nvidia GB200 systems.

Why it matters: Robot demonstration data is scarce and expensive to collect, but there’s plenty of video of the physical world. Developers of world-action models are betting that a model trained to predict video will need fewer robot examples. So far, the rankings favor that approach, but world action models tend to be larger and more demanding to run than VLAs. For developers, open weights and a fine-tuning recipe that works with a few hundred demonstrations mean a robotics research lab with a relatively low-cost arm and a powerful GPU can adapt the top-ranked model.

We’re thinking: FLUX 3 Action’s still fails more than half the time on simulated tasks. A simulated kitchen table is also tidier than a real one. We’d like to see FLUX 3 Action and its closest rivals on RoboArena next, to try their hands at more challenging robotics tasks.