German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio together. In BFL's early tests, it beat several rivals in video generation.
BFL describes Flux 3 as a step toward "real-world visual intelligence," which it defines as models that can "perceive, predict, and act across physical and digital environments." The company is part of a broader push to build so-called world models.
No single modality captures reality in full, BFL argues. Images show spatial structure, video captures how it changes over time, and audio can reveal links between mechanical events and the sounds they produce. Training on all three together lets them fill gaps for one another, giving the model more information than training on each modality separately.
Flux 3 adds native audio to videos up to 20 seconds long
Flux 3 can now generate videos with native audio for the first time, with clips up to 20 seconds long. It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. BFL says the model is especially good at human facial expressions and matching sounds to physical events.







