Black Forest Labs' FLUX 3 Beats Runway Gen-4.5 and Now Runs Factory Robots

Black Forest Labs launches FLUX 3, a single model that generates video, audio, images, and drives real robots on Audi's production floor

·
·
  • FLUX 3 launched: Black Forest Labs releases a single multimodal model jointly trained on video, audio, images, and action prediction. Blog post
  • Video in early access now: Generates up to 20-second, 720p videos with native audio; preferred over Runway Gen-4.5 in 77% of comparisons (self-reported).
  • Robots at Audi: FLUX-mimic, built on the FLUX 3 backbone with mimic robotics, is already deployed on Audi's production floor handling soft-body manipulation tasks.
  • 101ms reaction time: FLUX-mimic runs end-to-end in 101ms on a single RTX 5090, on par with human visual reaction time.
  • No pricing yet: No pricing disclosed for any tier; access via request form. FLUX 3 Image and open weights (FLUX 3 Dev) coming later this year.
  • Built on Self-Flow: A proprietary training technique that unifies generation and representation learning without external encoders like CLIP or DINO.

FLUX 3 is a single foundation model jointly trained across video, audio, and images from scratch, with the same backbone now running robots on an Audi factory floor. Black Forest Labs, the team behind the original FLUX image generation models, is betting that content creation and physical AI are the same problem.

One backbone, every modality

FLUX 3 jointly learns from images, video, and audio within a unified architecture. The core idea is that no single modality provides a complete picture of reality: each is a projection of the same underlying world, captured by different sensors. Images give you spatial structure. Video restores time and reveals physics. Audio exposes causal relationships that vision alone misses. The model reconciles all three at once, which forces it to build a coherent internal model of how the world actually works.

FLUX 3 builds on Self-Flow, BFL's approach for aligning multimodal generation and understanding within the same architecture. Self-Flow unifies representation learning and generation in one pass, without relying on frozen external encoders like CLIP or DINO. BFL then scaled compute and data substantially to train across all three modalities simultaneously. The result is a model that gets better at generating video because it understands the world better, not because it memorized more frames.

What it can do right now

FLUX 3 Video with optional native audio and FLUX 3 Action are available for early access. FLUX 3 Image will roll out in the coming weeks. No pricing has been disclosed as of this writing. You can request early access via the official form.

The video capability is the most developed. FLUX 3 generates videos up to 20 seconds long, with audio, in a single pass. Supported generation modes include:

  • Text-to-video and image-to-video generation

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves