Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio together. The headline feature is native audio generation: Flux 3 can create videos up to 20 seconds long with sound, a first for the company’s model family.

The model supports text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue, and linked clips for longer multi-shot sequences. BFL says the model is especially strong at facial expressions and matching sound to physical events.

The company positions Flux 3 as a step toward “real-world visual intelligence,” or models that can perceive, predict, and act across physical and digital environments. It has also developed Flux-mimic, a video action model for robotics tasks that is already being tested at Audi.

BFL’s own evaluations say Flux 3 beat several rivals in preference tests, but independent results are not yet available. For now, the release mainly shows how quickly generative media models are converging on video, sound, and action understanding in one system.