Black Forest Labs releases FLUX 3, a multimodal flow model generating 20-second video with native audio and actions

FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's…