Black Forest Labs' FLUX 3 Generates Video, Audio, and Images From One Model
Black Forest Labs launches FLUX 3 Video with native audio, up to 20-second clips, a draft-then-enhance workflow, and pricing that undercuts most competitors.
- FLUX 3 Video is now live via the BFL API, generating clips up to 20 seconds at 1080p with native synchronized audio in one request.
- Natively multimodal: audio, video, and images are trained together in one architecture (Self-Flow), not stitched from separate models, unlike Runway, Kling, and Luma.
- Draft mode lets you preview at $0.06/s, then deterministically upgrade the winner to full quality at $0.17/s HD, avoiding wasted full-render spend.
- Benchmark claims: BFL's internal tests show FLUX 3 preferred over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%), but near-even with Seedance 2.0 and Gemini Omni; numbers are vendor-reported and unverified.
- Multilingual lip-sync across 14+ languages is a standout feature; community testing shows it excels at dialogue-driven scenes but trails Seedance on high-energy action.
- Open weights (FLUX 3 Dev) and 2K/4K output are confirmed for later in 2026; the same backbone already powers a robotics model running on Audi factory floors.
FLUX 3 Video is now generally available via the BFL API and select partner tools. Black Forest Labs, the team behind the FLUX.1 image models, is making its first move into video generation with an architecture that differs from every other system in the field. Rather than training separate image, video, and audio models and wiring them together, BFL trained one model jointly across all three signals using their Self-Flow approach.
One model, not three
Self-Flow is BFL's method for aligning multimodal generation and understanding within a single architecture. FLUX 3 scales that approach with significantly more compute and data, training on video, images, and audio simultaneously. The design logic is straightforward: a model that jointly learns from all three modalities can build a representation of the world rather than a representation of any one format. It learns how objects hold together, how things move, and how events sound. No single modality provides a complete description on its own.
For audio, this has a concrete consequence. Joint training means audio is generated from the same flow-matching pass as the video frames, so a single generation call produces a complete audiovisual clip with no separate diffusion model for sound and no lip-sync post-processing. Every major competitor in the space, including Runway, Luma, and Kling, assembles audio and video through separate pipelines or attaches audio as a post-generation pass.
What it can actually do
The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video. The full capability set at launch:
- Text-to-Video: Simple or complex prompts, with natural movement, scene logic, and audio generated together.
- Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. FLUX 3 Video connects these in sequence while following the intended visual language.