Yann LeCun's AD-E2E-JEPA Plans Driving Trajectories 100x Faster Without Policy Training

A new JEPA world model drives without a trained policy, hitting goals 20 meters out with a 100x inference speedup over prior latent planners.

·
·
Yann LeCun's AD-E2E-JEPA Plans Driving Trajectories 100x Faster Without Policy TrainingPRO
  • AD-E2E-JEPA drives via goal-conditioned zero-shot planning, no policy training required.
  • A SIGReg-regularized projector cuts planning patches 16x and embedding dim 4x for a 100x speedup.
  • An 8-frame rollout over 256 trajectories runs in roughly 0.8 seconds.
  • Reaches 20-meter-away goals within 2.8 meters using 8,192 candidate trajectories.
  • Scores 67.3/72.9 EPDMS on NAVSIMv2 navtest in zero-shot planning mode.
  • Pretrained projector lifts downstream imitation learning from 80.2 to 85.4 EPDMS; code released.

AD-E2E-JEPA Plans Driving Trajectories Without a Learned Policy

A research team including Yann LeCun has proposed AD-E2E-JEPA, a self-supervised world model that selects driving trajectories without training an imitation policy. The AD-E2E-JEPA paper describes a planner that predicts candidate rollouts in latent space, compares them with a supplied goal observation, and chooses the closest match.

The model reportedly runs an eight-frame rollout over 256 trajectory candidates in about 0.8 seconds, roughly 100 times faster than earlier high-accuracy JEPA planners evaluated by the authors. Its compressed representation also improves a downstream imitation planner from 80.2 to 85.4 EPDMS on NAVSIMv2.

Most end-to-end driving systems train a policy to imitate human trajectories. That process combines visual representation, world prediction, and action selection inside one learned planner, which makes it difficult to isolate the source of a failure. AD-E2E-JEPA separates those components by evaluating the world model through goal-conditioned trajectory search.

At inference time, the planner receives the current sensor input, a future observation representing the goal, and a set of candidate trajectories. It then performs four steps:

  1. Encode the current observation and goal frame into latent representations.
  2. Roll each candidate trajectory forward for eight frames with the world model.
  3. Measure the distance between each predicted endpoint and the encoded goal.
  4. Select the candidate with the smallest latent-space distance.

Zero-shot planning refers to the action-selection stage: the system queries the pretrained world model directly and requires no separate policy-training phase. This setup gives researchers a cleaner measure of whether the model has learned representations useful for planning.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads