Black Forest Labs is steering its FLUX 3 foundation away from pixels and toward robot arms with FLUX 3 Action, a downloadable 7-billion-parameter system that reads camera feeds plus joint angles, outputs motor commands and forecasts how the scene should change as it goes.
Robotics developers usually face an uncomfortable choice. Policies that forecast upcoming frames together with movements win leaderboards yet demand heavy compute, while lighter vision language action systems react faster but botch more jobs. Black Forest Labs says its one-step checkpoint closes much of that gap, scoring 38.3% on the RoboLab-120 simulated benchmark versus 36.8% for NVIDIA's Cosmos 3 Nano and 28.0% for Pi0.5. A slower guidance-distilled version climbs to 42.2%, by its tally.
The throughput numbers deserve a closer read. The firm claims its FP8 checkpoints outpace Cosmos 3 Nano by 1.52x to 3.95x on gaming, pro and server graphics cards, and beat Pi0.5 by 1.34x to 2.28x on pro and server hardware. Yet its published tables concede a loss on the RTX 5090, where Pi0.5 comes out 1.15x to 2.42x quicker depending on numeric format. Those comparisons use real-time factor, not per-request delay, because every inference spans 2.13 seconds of arm movement against 1.0 second for Pi0.5.
Evidence from physical hardware is thinner. Positronic Robotics, in a trial Black Forest Labs calls independent, put the model on a Franka arm for ten chores with three tries apiece, and the tester was blind to which system was driving. FLUX 3 Action finished 28 of 30, edging Cosmos 3 Nano's 27 and well ahead of DreamZero's 20 and Pi0.5's 13. With so few trials, a one-try margin over NVIDIA's entry sits within noise.
The sharper case is about money. Black Forest Labs paired GPT 6 Astra, a large reasoning system, as a supervisor that hands routine motion to FLUX 3 Action and intervenes only when stuck. That combination cleared 90% of episodes at $8.77 and 8 minutes per win, the company says, versus $13.47 per win when Astra reasons alone at maximum effort, which remains the most accurate option. The logic is intuitive: a competent reflex layer lets the pricey planner stay quiet more of the time.
Most of what the model learned came from footage, with video above 95% of pretraining tokens, followed by a midtraining blend of gameplay clips, first-person hand recordings, handheld gripper captures and teleoperation spanning 14 robot bodies. Black Forest Labs positions games as a proving ground for navigation and eventual computer-use agents, and thanks NVIDIA for help bringing the model to Jetson, Hugging Face and LeRobot. Weights ship with a detailed technical write-up, but the page stops short of stating licensing conditions, something any company planning a deployment should confirm first.













