RLinf's RPent Wraps Frozen Pi0.5 Robots With an LLM Recovery Brain
RLinf open-sourced RPent, a recursive agent framework that wraps frozen vision-language-action models with planners and memory to reliably run manipulation tasks.
- RLinf released RPent, an Apache 2.0 agentic framework for embodied AI in the physical world.
- Wraps frozen VLAs like Pi0.5 as tools called by Claude Code or Codex planners.
- Ships benchmarks on LIBERO-PRO, LIBERO, RoboCasa365-Target50, and RoboTwin suites.
- Companion paper Harness VLA details memory-guided agent steering of frozen VLAs.
- Includes browser dashboard, interactive CLI, and endpoints for distributed env and VLA servers.
- Supports LIBERO and RoboCasa simulators plus Franka and SO-101 real-world robots.
RPent gives frozen VLAs an LLM control loop
Vision-language-action models such as Pi0.5 map camera observations and language instructions to robot actions, yet long tasks expose brittle planning and weak failure recovery. Robotics teams also tend to build separate evaluation harnesses around each model. RLinf has open-sourced the RPent repository, which places a frozen VLA behind an LLM planner such as Claude Code or Codex and connects it to perception, simulation, memory, and robot-control services.
RPent, short for Recursive Physical Agent, expands the agent’s context after every action. The planner observes the resulting state, checks progress, records useful outcomes, and chooses the next primitive. The VLA weights remain fixed during evaluation; adaptation happens through task decomposition, tool selection, and memory.
One model becomes a callable skill
RPent separates the stack into four layers connected through common interfaces. Components can run in one process or as services on different machines, allowing a simulator, VLA server, and planner to use separate compute.
| Layer | Responsibility | Current options |
|---|---|---|
| User | Launches runs and monitors episodes | Command-line interface and browser dashboard |
| Intelligence | Plans tasks, stores memory, and invokes action primitives | Claude Code, Codex, custom planners, Pi0.5, RLDX-1, and DreamZero |
| Interface | Provides a common robot-control boundary | Local processes, HTTP endpoints, and socket endpoints |
| Environment | Executes actions and returns observations | LIBERO-PRO, RoboCasa, Franka, and SO-101 |
A typical episode moves through the following control loop:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.