
FLUX 3 is not an image model with video bolted on. It is a single foundation model, jointly trained across video, audio, and images from scratch, with the same backbone now running robots on an Audi factory floor. Black Forest Labs, the team behind the original FLUX image generation models, is making a much larger bet: that content creation and physical AI are the same problem.
One model to rule all modalities
FLUX 3 jointly learns from images, videos, and audio within a unified architecture. The core idea is that no single modality provides a complete picture of reality -- each is a projection of the same underlying world, captured by different sensors. Images give you spatial structure. Video restores time and reveals physics. Audio exposes causal relationships that vision alone misses. The model is forced to reconcile all of them at once, which means it has to build a coherent internal model of how the world actually works.
FLUX 3 builds on Self-Flow, BFL's approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, they significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio simultaneously. Self-Flow is a training technique that unifies representation learning (understanding the world) and generation (producing outputs) in one pass, without relying on frozen external encoders like CLIP or DINO. The result is a model that gets better at generating video because it understands the world better, not just because it memorized more frames.
What it can actually do today
FLUX 3 Video with optional native audio generation and FLUX 3 Action are now available for early access, while FLUX 3 Image will roll out in the coming weeks. No pricing has been disclosed for any tier as of this writing. You can request early access via the official form.
The video capability is the most fleshed-out right now. FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation. The full list of supported generation modes includes:
- Text-to-video and image-to-video generation
- Video-to-video: carry a character or element from one scene into a new context
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves
