UW and NVIDIA's SGS Trains Real Robots Across a Million Simulated Environments

Success-Guided Sampling fixes a wasted-batch problem in massively parallel RL, letting PPO scale past a million simulated robots and transfer to real hardware.

·
·
·
UW and NVIDIA's SGS Trains Real Robots Across a Million Simulated EnvironmentsPRO
  • SGS is a sampler tweak on top of PPO that focuses training on tasks at the policy's capability frontier.
  • Scales cleanly to 2^20 parallel environments where uniform sampling and PLR collapse or stall.
  • Franka nut-and-bolt reaches 0.70 success at 1M envs vs 0.05-0.06 for baselines.
  • ANYmal D locomotion hits 0.73 across all terrains at 1M envs vs 0.54 for PLR.
  • UR5e assembles NIST board parts zero-shot from RGB, trained only in simulation.
  • Project site and paper live; code coming soon, CoRL 2026.

Success-guided sampling scales robot RL past one million environments

A University of Washington and NVIDIA team reports that task selection limits large-scale sim-to-real reinforcement learning more than the underlying training algorithm. Its method, Success-Guided Sampling (SGS), modifies the reset sampler so each simulated robot spends more time on configurations near its current ability. Paired with standard proximal policy optimization (PPO), SGS scaled beyond one million parallel environments.

In zero-shot hardware tests, a UR5e robot used RGB observations to thread nuts onto bolts, mesh gears, and insert pegs on a NIST assembly task board. Zero-shot means the policies trained entirely in simulation and received no real-world fine-tuning before deployment. The reported system also used no human demonstrations or per-task reward shaping. The same training recipe taught ANYmal quadrupeds to cross floating islands and stepping stones and to climb boxes.

SGS results across locomotion terrains and robot assembly tasks
SGS was evaluated on contact-rich assembly and quadruped locomotion tasks.

Why giant batches go stale

GPU simulators such as Isaac Lab can run thousands of environment copies and pool their trajectories into a single PPO update. A larger batch provides more experience, but its value depends on which tasks those environments attempt.

Uniform reset sampling assigns each worker a random task configuration, defined as an initial robot state, a goal, and an environment layout such as terrain or object pose. As training progresses, many sampled configurations become either routine or temporarily unreachable. Both groups contribute limited learning signal, while configurations near the policy’s competence boundary become a smaller share of the batch.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads