Black Forest Labs made history when they released their FLUX.1 image model. They were one of the first independent research labs to pioneer fast, high-quality, open source image generation. Now, they are taking a swing at video, and we are intrigued.

FLUX 3 is the lab’s new multimodal foundation model. Compared to other labs, they are taking a new approach. BFL has developed a model with a consolidated architecture, learning from images, audio, and video to generate multimodal outputs. Taken straight from their release blog:

“Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.

Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.”