• Sources: Black Forest Labs blog, HN 49033127
  • Summary: Black Forest Labs published FLUX-mimic on 2026-07-23, a video-action model built with mimic robotics on the FLUX 3 backbone announced the same week. A lightweight action decoder is trained on intermediate features from FLUX 3's video-prediction pathway, so actions are read out of the representation the video model already learned rather than from a separately trained control policy. The post states video prediction accounts for over 95% of training compute, that adding action prediction to the curriculum cost about 10% performance before recovering within 3,500 steps, and reports state-of-the-art success rates when fine-tuned plus roughly 2x sample efficiency against video-only models, with deployment on soft-part manipulation at Audi production facilities. Weights, API access, and license terms are not stated.
  • Comments: HN commenters read the result as a video generation model containing a usable world representation that transfers to control, and noted a demonstration where a robot arm needed three attempts to reseat window trim.
  • Why it matters: If action decoding rides on generic video pretraining, robotics data collection stops being the main constraint and video-model scale becomes the lever.
  • Follow-up: Watch for weights, license, and any evaluation that is not vendor-run.

send feedback on this story