RLinf's RPent Wraps Frozen Pi0.5 Robots With an LLM Recovery Brain

RLinf open-sourced RPent, a recursive agent framework that wraps frozen vision-language-action models with planners and memory to reliably run manipulation tasks.

·
·
RLinf's RPent Wraps Frozen Pi0.5 Robots With an LLM Recovery BrainPRO
Read2 min
TypeRepo
  • RLinf released RPent, an Apache 2.0 agentic framework for embodied AI in the physical world.
  • Wraps frozen VLAs like Pi0.5 as tools called by Claude Code or Codex planners.
  • Ships benchmarks on LIBERO-PRO, LIBERO, RoboCasa365-Target50, and RoboTwin suites.
  • Companion paper Harness VLA details memory-guided agent steering of frozen VLAs.
  • Includes browser dashboard, interactive CLI, and endpoints for distributed env and VLA servers.
  • Supports LIBERO and RoboCasa simulators plus Franka and SO-101 real-world robots.

RPent gives frozen VLAs an LLM control loop

Vision-language-action models such as Pi0.5 map camera observations and language instructions to robot actions, yet long tasks expose brittle planning and weak failure recovery. Robotics teams also tend to build separate evaluation harnesses around each model. RLinf has open-sourced the RPent repository, which places a frozen VLA behind an LLM planner such as Claude Code or Codex and connects it to perception, simulation, memory, and robot-control services.

RPent, short for Recursive Physical Agent, expands the agent’s context after every action. The planner observes the resulting state, checks progress, records useful outcomes, and chooses the next primitive. The VLA weights remain fixed during evaluation; adaptation happens through task decomposition, tool selection, and memory.

Diagram of RPent's user, intelligence, interface, and environment layers
RPent separates planning, action primitives, robot interfaces, and environments into composable services.

One model becomes a callable skill

RPent separates the stack into four layers connected through common interfaces. Components can run in one process or as services on different machines, allowing a simulator, VLA server, and planner to use separate compute.

Layer Responsibility Current options
User Launches runs and monitors episodes Command-line interface and browser dashboard
Intelligence Plans tasks, stores memory, and invokes action primitives Claude Code, Codex, custom planners, Pi0.5, RLDX-1, and DreamZero
Interface Provides a common robot-control boundary Local processes, HTTP endpoints, and socket endpoints
Environment Executes actions and returns observations LIBERO-PRO, RoboCasa, Franka, and SO-101

A typical episode moves through the following control loop:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads