Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.

Black Forest Labs launches FLUX 3, a multimodal AI model generating 20-second videos with synced audio and powering robotics on Audi production

FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.