RAI Institute's Exploy Ships Robot AI Policies in a Single Portable File
RAI Institute's new open-source library packages entire RL control pipelines into one ONNX graph, eliminating the C++ rewrites that plague sim-to-real transfer.
- RAI Institute released Exploy, an open-source library for deploying RL policies to robots.
- Packages observation, network, and action logic into one self-contained ONNX computational graph.
- Eliminates manual C++ rewrites and fixes Python-to-C++ numerical discrepancies that cause unsafe behaviors.
- Includes adapters for IsaacLab and MjLab, supports MLPs, RNNs, CNNs, Transformers, and diffusion policies.
- Ships with a C++ runtime on ONNX Runtime, plain CMake libraries, and a ROS 2 package.
- Already deployed on Spot, Unitree G1, Atlas, Roadrunner, UMV, and an autonomous bike.
Exploy puts a robot policy into one ONNX file
The RAI Institute has released Exploy, an open-source library that exports a reinforcement learning policy’s full control path into one ONNX artifact. It targets a costly step in simulation-to-robot deployment: reproducing Python observation and action logic inside a robot’s C++ control stack.
The missing half of ONNX export
ONNX, short for Open Neural Network Exchange, provides a portable format for running trained models across frameworks and hardware. Most robotics workflows export the policy network’s graph and weights, leaving sensor processing, observation construction, recurrent state, action scaling, and command generation in the training environment.
Policy behavior depends on those surrounding operations. Developers must recreate them in C++, where changes in constants, operation order, normalization, or floating-point behavior can alter the commands sent to hardware.
That handoff creates three recurring problems:
- Numerical drift: Python and C++ implementations can produce different outputs from the same sensor data, potentially generating unsafe commands.
- Iteration cost: Each policy or environment change can require another round of integration work.
- Codebase divergence: Training environments and robot controllers evolve separately, making results harder to reproduce.
One graph from sensors to commands
Exploy expands the exported boundary to include the PyTorch operations around the policy network. The resulting .onnx file can map raw sensor measurements to low-level actuator commands or higher-level targets such as desired velocity.
The export and deployment process has four stages:
- Register the interface. A unified context manager records tensor inputs and outputs, persistent state for recurrent models, and metadata such as control rates, joint stiffness, and damping gains.
- Trace the pipeline. The exporter captures observation generation, the neural network forward pass, and action post-processing from the source environment.
- Check numerical parity. Evaluation tools compare the original PyTorch pipeline with the exported ONNX artifact step by step before hardware deployment.
- Run the controller. A lightweight C++ runtime built on ONNX Runtime discovers the artifact’s tensor interfaces and connects them to the robot’s hardware state and command interfaces.
Packaging the graph and configuration together allows the same C++ deployment stack to load different compatible policies without per-policy rewrites. Each robot must still provide the sensor and command interfaces expected by the artifact.
Where Exploy fits
| Area | Support |
|---|---|
| Training frameworks | Built-in adapters for IsaacLab and MjLab, plus an extensible interface for custom PyTorch environments |
| Model architectures | MLPs, CNNs, LSTMs, GRUs, transformers, diffusion policies, and other models whose operations ONNX can represent |
| Native deployment | CMake libraries and a ROS 2 package |
| Runtime | A C++ controller using ONNX Runtime |
From Roadrunner to Atlas
RAI reports using Exploy across wheeled, legged, balancing, and manipulation platforms:
- Roadrunner, RAI’s custom robot, uses one reinforcement learning policy to switch among Segway-style balancing, bicycle driving, and stepping. The policy was transferred from simulation without hardware fine-tuning.
- Boston Dynamics’ Spot runs depth-based and recurrent policies for high-obstacle traversal and parkour courses.
- Unitree G1 humanoids and Boston Dynamics’ Atlas use exported policies for navigation, blind walking, and rough-terrain traversal.
- RAI’s autonomous bicycle uses Exploy for trail racing, while the UMV and AthenaZero platforms use it for other mobility and manipulation work.
The remaining sim-to-real work
ONNX operator coverage defines Exploy’s export ceiling, so operations that ONNX cannot represent fall outside the automatic path. Hardware integrations also need compatible sensors, control rates, state interfaces, and command semantics.
Numerical equivalence tests verify parity between the PyTorch and ONNX implementations. Simulation fidelity, sensor calibration, communication latency, actuator dynamics, fault handling, and hardware safety validation remain separate engineering tasks.
The code is available in the Exploy repository.