NYU's H-JEPA Boosts Robot Planning Success from 18% to 73%
A hierarchical JEPA trained end-to-end lifts Visual AntMaze planning success from 18% to 73% while using less planner compute.
- H-JEPA stacks action-conditioned JEPA world models, each predicting farther ahead in its own learned latent.
- On Visual AntMaze, a three-level hierarchy raises success from 18% to 73% at lower planner compute.
- Higher levels automatically drop fast, unpredictable details and keep slow task-relevant state like agent position.
- Planning is top-down: each level's predictions become subgoals for the level below it.
- Benefit only appears when training data shows at least a 2x gap between fast and slow components.
- Code released at github.com/kevinghst/H-JEPA, built on LeWM and stable-worldmodel.
H-JEPA plans farther by learning at multiple timescales
Long-horizon planning from pixels forces a world model to preserve fine motor detail while tracking slow progress toward a distant goal. A new H-JEPA preprint from researchers at NYU, Meta’s AMI Labs, INRIA Paris, and Brown University proposes a hierarchy of action-conditioned world models, each with its own latent representation and temporal stride.
H-JEPA plans from coarse goals down to concrete actions. The approach improves success across the paper’s simulated navigation and manipulation tasks while reducing planner computation, provided the underlying data contains clearly separated timescales.
One latent space, two conflicting jobs
A joint-embedding predictive architecture, or JEPA, encodes observations into compact latent representations and predicts future latents conditioned on actions. Planning then searches for actions whose predicted outcome approaches the encoded goal.
Long rollouts expose a conflict in flat models. Low-level dynamics require details such as joint angles and object orientation, while goal matching may depend mainly on slower variables such as maze position. Packing both into one representation can make the planning cost sensitive to details that have little bearing on reaching the goal.
Planning expense also rises quickly with horizon length. Every additional step expands the action sequence that the optimizer must search, making distant goals expensive even when the route can be described with a few coarse decisions.
Longer horizons produce coarser state
H-JEPA stacks several JEPA modules at progressively longer temporal strides. The lowest level predicts frame-to-frame changes, intermediate levels model wider intervals, and the highest level may span much of an episode. Gradients update the full stack through one end-to-end training objective.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.