NVIDIA's Physis-Lang Teaches Video AI Real Physics, Beats Google's Veo 3.1

NVIDIA's Physis-Lang teaches video world models to reason about physics through evolving captions, lifting Cosmos 3 to the top of Physics-IQ.

·
·
  • NVIDIA released Physis-Lang, a self-evolving framework adding physics reasoning to video captions and prompts.
  • Cosmos 3 Super and Nano with Physis-Lang take rank 1 and 2 on Physics-IQ Verified at 48.2% and 43.3%.
  • Adding physics reasoning to prompts alone lifted Cosmos 3's PhyGenBench score by 5.62 points with no retraining.
  • A physics-aware critic evolves captioning guidelines; failures drive language-guided retrieval of missing physical phenomena.
  • New PhysCapBench evaluates caption quality across 246 videos and 3,794 human-verified atomic assertions.
  • Gains generalize across Wan 2.1 14B and all Cosmos 3 sizes, from 4B Edge to 64B Super.

NVIDIA’s Physis-Lang teaches video models through physics-aware captions

NVIDIA researchers introduced Physis-Lang, an open framework that encodes physical explanations in the captions and prompts used by video-generation models. The work targets a persistent weakness in generated video: realistic frames can still show butter melting upward, blocks passing through one another, or cloth moving without plausible weight and tension.

Training captions provide supervision about objects, actions, and temporal changes. Physis-Lang enriches that supervision with cause, governing principle, and effect. A melting-butter caption, for example, describes heat raising the temperature, the material changing phase, gravity deforming the softened mass, and a liquid pool forming. The framework uses this language to recaption training videos and expand user prompts before generation.

Better prompts lift physics scores

In the reported Physics-IQ Verified snapshot, Physis-Lang configurations of Cosmos 3 Super and Cosmos 3 Nano occupy the top two positions. The benchmark evaluates how faithfully generated videos follow real-world physical behavior.

Reported Physics-IQ Verified results
Configuration Score
Physis-Lang with Cosmos 3 Super 48.2%
Physis-Lang with Cosmos 3 Nano 43.3%
Cosmos 3 Super baseline 42.7%
Seedance 2.5 42.4%
Physics-IQ leaderboard scores before and after applying Physis-Lang
Physis-Lang configurations lead the reported Physics-IQ Verified comparison.

An inference-only ablation isolates the effect of prompt expansion. Adding physics-aware reasoning to the input increased Cosmos 3’s PhyGenBench score by 5.62 points with fixed model weights and no additional training. This path can be applied to an existing checkpoint because it changes the conditioning text sent to the generator.

A critic rewrites the captioning rules

The framework shares one physical-language representation across three connected stages, allowing critic feedback and failure-driven retrieval to improve the text used for training and generation.

  1. Refine the language. A physics-aware critic checks each caption for missing claims and statements unsupported by the video. An agent updates the captioning guidelines, while the captioner remains frozen during this loop.
  2. Retrieve weak phenomena. Failed generations become search queries. Language-guided retrieval selects visually diverse videos covering the physical categories where the system performs poorly.
  3. Train and generate. The revised guidelines recaption training data and expand inference prompts. Scene-specific negative descriptions identify implausible dynamics for the generator to suppress.
Diagram of the Physis-Lang caption refinement, retrieval, training, and generation pipeline
The feedback loop updates captioning guidelines, retrieves relevant videos, and reuses the resulting language throughout the pipeline.

PhysCapBench measures whether the captions improve independently of video generation. The benchmark contains 246 reviewed videos and 3,794 human-verified physical assertions, averaging 15.4 assertions per video. Evaluators split each generated caption into atomic claims, check those claims against the source video, and calculate precision, recall, and F1, which balances the two.

Results across nine caption-guideline iterations
Metric Initial score Final score
PhysCapBench caption F1 78.64 87.82
Downstream PhyGenBench 64.17 67.29

Failures choose the next training clips

On VideoPhy-2 Hard, targeted retrieval raised joint scores in all eight reported categories. Gains ranged from 1.23 points for elasticity to 8.00 points for thermal and chemical processes.

VideoPhy-2 Hard gains by physical category
Category Joint-score gain
Thermal processes +8.00
Chemical processes +8.00
Fracture mechanics +7.45
Soft-body motion +7.41
Cloth deformation +7.19
Rigid-body motion +6.06
Contact and collision +3.45
Elasticity +1.23

The largest gains occur in processes governed by partially hidden state variables such as temperature, material phase, and chemical composition. Explicit causal descriptions provide conditioning signals for transitions that individual frames may not reveal directly.

Gains carry across backbones

Table 5 of the release paper reports positive mean gains for every tested backbone, spanning models from 4 billion to 64 billion parameters.

Reported mean gain by video-model backbone
Backbone Parameters Mean gain
Wan 2.1 14B +7.05
Cosmos 3 Edge 4B +3.24
Cosmos 3 Nano 16B +6.22
Cosmos 3 Super 64B +5.02

Physis-Lang with Cosmos 3 Nano exceeds Veo 3.1 on PhyGenBench but falls slightly short on VideoPhy-2 All.

Aggregate benchmark comparison
System PhyGenBench VideoPhy-2 Hard
Physis-Lang with Cosmos 3 Nano 71.04 62.36
Veo 3.1 65.63 58.43

The benchmarks set clear boundaries

The 5.62-point ablation measures prompt expansion with fixed weights. The full pipeline also uses recaptioned data and failure-driven retrieval, so its results reflect the combined effect of language refinement, data selection, training, and inference-time prompting.

The quantitative evidence covers selected benchmark phenomena, and the top Physics-IQ Verified score remains 48.2%. General simulation ability, long-horizon consistency, and performance on unseen physical processes remain open evaluation questions. Held-out phenomena and human review would help determine how well the gains transfer beyond the reported test sets.

Captions become a control surface

Physics-aware captions give developers a model-agnostic way to supply causal structure. A caption that names forces, state changes, and expected outcomes provides more temporal guidance than a surface description of visible objects and actions. The approach operates at the data and prompt layers, allowing it to sit on top of existing checkpoints.

An implementation can begin with inference-time prompt expansion: define a domain-specific causal schema, expand incoming prompts, keep the generator fixed, and evaluate on a held-out suite. Training integration adds critic-guided recaptioning, retrieval of failure cases, and another model-training run. The same workflow can be tested in other conditional generative systems that use scene descriptions.

The project page and paper provide the framework details, benchmark definitions, ablations, and generated-video galleries covering collisions, tearing, wetting, and deformation under load.

Trending
  • No trending articles

Comments

avatar

Next Reads