Meta AI's NAVA-WAM Teaches Robots From Unlabeled Video, Hitting 93% Success
A new world action model skips latent action detours and learns robot control priors straight from raw video, posting 73.6% on RoboTwin Random.
- NAVA-WAM pretrains a robot action policy directly on observation-only video, skipping latent-action detours.
- Two-stage training: future-video flow matching, then joint video-action post-training with asymmetric attention.
- Hits 88.5% on RoboTwin 2.0 Clean and 73.6% on the Random OOD split, +17.9 over the best baseline.
- Real Franka FR3 deployment: 93.3% average success across 15 trials, beating π0.5 and DreamZero.
- Inference costs 15.8 ms per executed action on an H100, thanks to action-only denoising at test time.
- Project page with videos and benchmarks at zhaochongan.github.io/projects/NAVA-WAM.
NAVA-WAM Pretrains Robot Policies on Unlabeled Video
Researchers from Meta AI, the University of Copenhagen, Imperial College London, and Physical Intelligence have developed a way to train robot policies from observation-only video. According to the paper, NAVA-WAM improves out-of-distribution performance in simulation and completes 14 of 15 real-world trials on a Franka arm.
Observation-only video shows how objects and people move but omits the synchronized robot commands required for imitation learning. NAVA-WAM routes a video-prediction objective through the action policy, then grounds that policy with action-labeled robot demonstrations. This design lets passive video shape the controller itself.
Action labels cap the data supply
Robot policies typically map camera observations and instructions to chunks of motor commands. Collecting those commands requires teleoperation or autonomous rollouts on physical robots, making action-labeled data slower and more expensive to acquire than ordinary video.
Existing video-pretraining methods generally use one of three designs:
- Vision representation learning: A video encoder learns visual features, which a separately trained policy consumes.
- Latent action learning: A model infers pseudo-actions from visual transitions, then maps those representations to robot commands.
- Native action-prior learning: NAVA-WAM sends the video-prediction signal through the action model during pretraining, removing the intermediate feature or pseudo-action interface.
Routing video through action tokens
The architecture pairs two Diffusion Transformers: a Video-DiT for visual states and an Action-DiT for robot commands. Both use flow matching, a diffusion-style objective that trains a model to transform noise into structured samples by predicting the direction of that transformation.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.