Black Forest Labs' FLUX 3 Generates Video, Audio, and Images From One Model
Black Forest Labs launches FLUX 3 Video with native audio, up to 20-second clips, a draft-then-enhance workflow, and pricing that undercuts most competitors.
- FLUX 3 Video is now live via the BFL API, generating clips up to 20 seconds at 1080p with native synchronized audio in one request.
- Natively multimodal: audio, video, and images are trained together in one architecture (Self-Flow), not stitched from separate models, unlike Runway, Kling, and Luma.
- Draft mode lets you preview at $0.06/s, then deterministically upgrade the winner to full quality at $0.17/s HD, avoiding wasted full-render spend.
- Benchmark claims: BFL's internal tests show FLUX 3 preferred over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%), but near-even with Seedance 2.0 and Gemini Omni; numbers are vendor-reported and unverified.
- Multilingual lip-sync across 14+ languages is a standout feature; community testing shows it excels at dialogue-driven scenes but trails Seedance on high-energy action.
- Open weights (FLUX 3 Dev) and 2K/4K output are confirmed for later in 2026; the same backbone already powers a robotics model running on Audi factory floors.
FLUX 3 Video is now generally available via the BFL API and select partner tools. Black Forest Labs, the team behind the FLUX.1 image models, is making its first move into video generation, and the approach is architecturally different from everything else in the field. This is not a video model bolted onto an image model. Rather than training an image model, a video model, and an audio model and wiring them together, BFL trained one model jointly across all of those signals within one architecture built on their Self-Flow approach.
One model, not three
FLUX 3 builds on Self-Flow, BFL's approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this, they significantly scaled up compute and data to train FLUX 3 across video, images, and audio at the same time. The key insight driving the design: the model jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, it must learn a representation of the world: how objects hold together, how things move, and how events sound. No single modality provides a complete description.
This has a concrete consequence for audio. Joint multimodal training means audio is generated from the same flow-matching pass as the video frames. No separate diffusion model for sound. No lip-sync post-processing. A single generation call produces a complete audiovisual clip. Every major competitor in the video-generation space, including Runway, Luma, and Kling, assembles audio and video through separate pipelines or attaches audio as a post-generation pass.
What it can actually do
The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video. The full capability set at launch:
- Text-to-Video: Simple or complex prompts, with natural movement, scene logic, and audio generated together.
- Image-to-Video and Keyframes: