FLUX 3 lands: one model for image, video, and audio
Black Forest Labs' FLUX 3 generates 20-sec video with native audio from a single unified multimodal model.

Black Forest Labs just dropped FLUX 3, a multimodal foundation model that treats image, video, and audio as one problem instead of three. FLUX 3 Video spins up clips up to 20 seconds long with native audio from text, image, or video prompts — plus video continuation, keyframe transitions, multilingual dialogue, and clip chaining for longer sequences.
The trick is a unified architecture that learns spatial structure, motion, sound, and physical interaction together, built on the company's Self Flow method for aligning generation and understanding in the same system. There's even a FLUX 3 Action variant aimed at robotic action prediction.
For now it's early access only. BFL says API access, private weights, and an open-weight FLUX 3 Dev are all coming later this year.
Why it matters: the maker of the open-weight FLUX image models is now going head-to-head with Sora and Veo — and promising to open-source part of it.
Sources
Independent coverage
- Black Forest Labs Unveils FLUX 3, A New Multimodal Frontier Model For Visual Intelligence globenewswire.com
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio venturebeat.com
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.