Get the app
Industry

FLUX 3 lands: one model for image, video, and audio

Black Forest Labs' FLUX 3 generates 20-sec video with native audio from a single unified multimodal model.

FLUX 3 lands: one model for image, video, and audio

Black Forest Labs just dropped FLUX 3, a multimodal foundation model that treats image, video, and audio as one problem instead of three. FLUX 3 Video spins up clips up to 20 seconds long with native audio from text, image, or video prompts — plus video continuation, keyframe transitions, multilingual dialogue, and clip chaining for longer sequences.

The trick is a unified architecture that learns spatial structure, motion, sound, and physical interaction together, built on the company's Self Flow method for aligning generation and understanding in the same system. There's even a FLUX 3 Action variant aimed at robotic action prediction.

For now it's early access only. BFL says API access, private weights, and an open-weight FLUX 3 Dev are all coming later this year.

Why it matters: the maker of the open-weight FLUX image models is now going head-to-head with Sora and Veo — and promising to open-source part of it.

Sources

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play