Mila's RoboJEPA Lets Robotics Teams Predict Performance Before Buying Compute

A new paper fits a scaling law to robotic world models, trains an 8B JEPA on 15,000 hours of robot video, and plans on real hardware.

·
·
·
Mila's RoboJEPA Lets Robotics Teams Predict Performance Before Buying ComputePRO
  • New paper RoboJEPA establishes scaling laws for multi-embodiment robotic world models trained on real robot video.
  • Imagination error follows a second-order power law in compute, extrapolating accurately from 2B to 8B parameters.
  • Trained on 23 datasets, 12 embodiments, roughly 15,000 hours of video with 6,700 hours action-labeled.
  • Skills emerge in order with compute: arm motion, grasping, obstacle avoidance, then object pushing past 10^22 FLOPs.
  • On a real Franka, the 8B model grasps 67% zero-shot using only a goal image for conditioning.
  • All checkpoints, training code, and deployment code are released publicly.

RoboJEPA maps compute to robot world-model performance

Scaling laws let language-model teams estimate how additional compute will affect performance before committing to a training run. The RoboJEPA paper applies that approach to robotic world models, fitting a power law across predictors ranging from 22 million to 8 billion parameters.

The study connects three measurements that robotics researchers rarely obtain in one experiment: training compute, latent prediction error, and success on physical hardware. Its results suggest that teams can use an offline prediction metric to forecast some planning gains before spending time on robot evaluations.

RoboJEPA builds on Meta’s V-JEPA models, which learn visual representations from video without requiring manual labels for every frame. The author list includes researchers affiliated with Mila and collaborating institutions. The paper describes the 8B version as the largest JEPA predictor trained to date.

A forecast for world-model training

Latent world models compress observations into numerical representations called embeddings, then predict how those representations will change. A planner can test candidate action sequences inside this learned model and select the sequence whose predicted outcome best matches a goal.

This approach reduces the cost of predicting every pixel, but it leaves teams with a budgeting problem. Large training runs consume substantial compute, while physical evaluation requires robots, operators, reset procedures, and repeated trials. Without a scaling curve, developers have little evidence for estimating the return from a larger predictor.

A controlled test of predictor scale

RoboJEPA uses a frozen V-JEPA 2.1-G encoder to convert video frames into embeddings. The researchers scale only the transformer that predicts future embeddings, helping isolate the relationship between predictor capacity, training compute, and rollout accuracy.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads