Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
Black Forest Labs has released Flux 3, a multimodal foundation model that can generate video with native sound for the first time. The model produces clips up to 20 seconds long and, according to the company's internal tests, slightly outperforms Seedance 2.0. Black Forest Labs is also testing Flux 3 on robotics tasks as part of its broader goal to build a world model.
Why it matters: Flux 3 represents a notable advance in unified multimodal generation by adding native audio to video, potentially accelerating applications in content creation, simulation, and robotics.
Full story at: The Decoder ↗