Black Forest Labs' FLUX 3 Beats Sora 2 and Veo 3.1 on Physics
Black Forest Labs' new multimodal model tops Google DeepMind's physics benchmark for video generation, beating Sora 2, Veo 3.1 and Cosmos 3.
- FLUX 3 tops Physics-IQ Verified video-to-video leaderboard with +14.68 pp above track mean.
- Beats NVIDIA Cosmos 3, Seedance 2.5, MiniMax H3, Gemini Omni, Veo 3.1 and Sora 2.
- Base FLUX 3 scores 51.11; best-of-8 with physics verifier reaches 64.35.
- Benchmark from Google DeepMind tests fluid, optics, solids, magnetism and thermodynamics.
- Sora 2 lands near bottom at -9.93 pp despite its public profile.
- Verifier uses V-JEPA 2 surprise plus candidate consensus for minimum Bayes risk selection.
FLUX 3 takes the Physics-IQ lead with verifier-assisted generation
Black Forest Labs’ FLUX 3 [large] ranks first in the verified video-to-video view of the Physics-IQ leaderboard. Its system predicts how filmed experiments continue, then uses a physics-aware selector to choose among eight generated candidates. The reported result places it ahead of NVIDIA’s Cosmos 3 systems, ByteDance’s Seedance 2.5, Google’s Veo 3.1, and OpenAI’s Sora 2.
Real footage supplies the reference
Google DeepMind’s Physics-IQ benchmark gives a video model the opening frames and description of a physical experiment, then compares its continuation with recorded footage. The clips cover fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics. Models must predict effects such as pouring water, moving shadows, bouncing objects, and interacting magnets.
Scoring measures spatial accuracy, spatiotemporal alignment, and mean squared error. Spatial metrics test whether objects and effects appear in the expected locations. Spatiotemporal alignment measures whether events unfold in the right place and sequence. Mean squared error penalizes differences from the recorded continuation, with larger deviations receiving heavier penalties.
Aesthetic preference tests dominate video-generation evaluations, often rewarding visual polish or prompt adherence. Physics-IQ adds a recorded reference continuation, which constrains scoring around observed motion and interactions. The footage represents one realized outcome rather than every physically possible outcome, so the benchmark measures reconstruction accuracy more directly than broad causal understanding.
FLUX 3 leads by 2.88 points
In the reported leaderboard snapshot, FLUX 3 [large] finishes 14.68 percentage points above the track mean. NVIDIA’s Physis-Lang, built on Cosmos3 Super, follows at 11.80 points above the mean. The 2.88-point gap separates the top two entries.
| Model | Improvement over mean | Normalized cost per video |
|---|---|---|
| FLUX 3 [large] (Black Forest Labs) | +14.68 pp | $0.868 |
| Physis-Lang Cosmos3 Super (NVIDIA) | +11.80 pp | $0.823 |
| Cosmos3 Super (NVIDIA) | +6.25 pp | Not listed |
| Seedance 2.5 (ByteDance) | +6.00 pp | $2.838 |
| MiniMax H3 | +3.36 pp | $0.438 |
| Gemini Omni 1.1 Flash (Google) | −1.09 pp | $0.609 |
| Veo 3.1 Lite (Google) | −4.60 pp | $0.300 |
| Sora 2 (OpenAI) | −9.93 pp | $0.640 |
The net-improvement values are relative to the track mean, so they can change as models enter or leave the leaderboard even when their raw scores remain fixed. The normalized costs are the leaderboard’s reported estimates and may differ from production pricing or total application costs.
Across the component metrics, FLUX 3 ranks first in spatial accuracy, weighted spatial accuracy, and mean squared error. It ranks second in spatiotemporal alignment, where Physis-Lang takes the lead. Seedance 2.5 trails FLUX 3 by 8.68 points while carrying about 3.3 times the reported normalized cost. Sora 2 finishes 9.93 points below the track mean.
Eight generations feed one result
The winning entry measures a model-and-selector pipeline. Black Forest Labs reports a raw FLUX 3 score of 51.11 ± 0.44 without candidate selection. Its best-of-eight configuration reaches 64.35 ± 0.22, an increase of 13.24 points.
For each prompt, the system generates eight candidates from separate seed pools. It ranks them using WMReward, a V-JEPA 2 surprise score, and a consensus calculation that compares each candidate with the others using Physics-IQ metrics. The candidate with the best combined rank becomes the submitted continuation.
This selection process resembles minimum Bayes risk decoding, an inference method that chooses the candidate most consistent with a scoring model or candidate set. Here, the learned reward and surprise signals act as a physics critic, while the consensus term favors continuations that agree with the broader sample.
Best-of-eight inference adds generation and selection overhead that a raw model score does not capture. Meaningful deployment estimates therefore depend on whether the reported normalized cost includes all eight candidates, reward-model evaluation, consensus scoring, and discarded outputs.
Prompt rewrites move the score
A follow-up submission replaced Qwen3-VL-generated clip descriptions with rewrites from Claude Opus 5.5. Black Forest Labs reported a gain of roughly seven points from that change. The result shows that Physics-IQ scores depend on prompt construction alongside model weights, sampling strategy, and candidate selection.
Developers comparing systems should keep prompt-generation methods consistent or report them as part of the evaluated pipeline. A score produced with rewritten descriptions and inference-time search does not describe the base generator alone.
Physics fits BFL’s robotics push
Black Forest Labs describes FLUX 3 as its first natively multimodal architecture for generating images, video, and audio. The same family also powers FLUX 3 Action, a 7-billion-parameter open-weights world-action model intended for robotics.
A shared backbone must represent appearance, motion, object interaction, and action consequences across media. Short-horizon physical consistency supports commercial video generation by reducing implausible motion, and it supports robotics research by improving predictions about how a scene may change after an action.
Robotics systems require capabilities beyond video continuation, including long-horizon planning, control accuracy, uncertainty handling, and safety validation. Physics-IQ provides evidence about near-term visual prediction under recorded conditions; it does not establish those broader capabilities.
The practical readout
- Compare complete pipelines. FLUX 3’s leading result includes eight generations, learned scoring, consensus ranking, and candidate selection.
- Separate raw and selected scores. The base model records 51.11, while verifier-assisted selection raises the result to 64.35.
- Account for all inference costs. Best-of-eight generation can change latency and compute requirements even when normalized leaderboard pricing appears competitive.
- Control prompt construction. A roughly seven-point gain from rewritten descriptions makes prompting a material benchmark variable.
- Match metrics to the application. FLUX 3 leads overall, while Physis-Lang leads the spatiotemporal component.
Teams evaluating video models as simulators can use Physics-IQ alongside aesthetic, prompt-adherence, latency, and cost tests. The current results favor a system that combines a capable generator with inference-time search and physics-aware selection, while the large gap between its raw and selected scores shows how much of the lead comes from the surrounding pipeline.