USC and NVIDIA's BCG Trains Humanoid Robots to Run and Fight Without Falls

A new smoothing trick lets differentiable simulators keep stiff, realistic contact physics while still producing gradients stable enough to train humanoid policies that transfer zero-shot to a real Unitree G1.

·
·
USC and NVIDIA's BCG Trains Humanoid Robots to Run and Fight Without FallsPRO
Read2 min
TypePaper
SubtopicHumanoid · Manipulation · Rl
  • BCG stabilizes differentiable simulation under stiff contact by averaging gradients over perturbed rollouts at contact events.
  • Enables zero-shot transfer of Run, Jump, Fight, Dance policies to a real Unitree G1.
  • Trained with SHAC + ADD imitation reward in NVIDIA Warp on a single RTX 2080 Ti.
  • At κ=300, 0/5 falls in MuJoCo transfer vs 5/5 with soft contact (κ=50).
  • Beats PPO on 3 of 4 motions with only 64 parallel environments vs 4096.
  • Limitations: manual tuning of bundle parameters, extra memory/compute, tested only on foot-ground contact.

Bundled gradients help humanoid policies survive stiff contact

Researchers from USC, ETH Zürich, NVIDIA, and ZHAW have introduced Bundled Contact Gradients (BCG), a method for training robot policies with stiff contact inside a differentiable simulator. The team trained a Unitree G1 humanoid to run, jump, fight, and dance, then deployed the policies on the physical robot without hardware fine-tuning. BCG addresses a persistent sim-to-real problem: soft contact produces usable gradients but unrealistic motion, while stiff contact better matches hardware and destabilizes gradient-based optimization.

Stiff contact scrambles the gradient

Differentiable physics engines expose derivatives of simulated dynamics, allowing a trainer to backpropagate through a rollout and update the policy directly. This approach can require fewer simulator samples and less training time than model-free reinforcement learning. Many differentiable simulators smooth their contact models to keep those derivatives stable, which allows feet and other bodies to penetrate surfaces more deeply than they would under rigid contact.

Stiff contact behaves like a step function because contact force turns on sharply as two surfaces meet. Near that boundary, tiny changes in position or velocity can produce large derivative changes. Neighboring rollouts may therefore point the optimizer in conflicting directions, causing first-order policy updates to stall.

The paper measures the fidelity cost through foot penetration relative to MuJoCo. At contact stiffness κ=50, the mean discrepancy is 13.5 mm. Raising κ to 300 reduces it to 2.9 mm. Higher κ values produce stiffer contact and a closer approximation to the dynamics expected during transfer.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads