Black Forest Labs has unveiled FLUX 3, a multimodal foundation model trained jointly on images, video, and audio to support content generation and robotic action prediction.
The model uses a unified architecture designed to learn spatial structure, movement, sound, and physical interactions together rather than treating each format as a separate task. It builds on Self Flow, the company’s method for aligning multimodal generation and understanding within the same system.
FLUX 3 Video can generate clips with native audio lasting up to 20 seconds from text, images, or existing videos. It also supports video continuation, keyframe transitions, multilingual dialogue, typography, and the chaining of clips into longer sequences.
In preliminary tests conducted by Black Forest Labs, FLUX 3 was preferred over Grok Imagine Video in as many as 69% of comparisons, Kling v3 Pro in 60%, Runway Gen 4.5 in 77%, and Luma Ray 3.2 in 93%.
The company cautioned that the model and evaluation system remain in development and said full benchmark results and methodology will be published with broader availability.








