Microsoft's Rho Cuts Robot Training Data With a 5B Open-Weights Model
Microsoft Research releases an open-weights 5B vision-language-action model family with embodiment-specific checkpoints and a lightweight online adaptation module.
- Microsoft Research releases Rho, an open-weights 5B VLA family for bimanual manipulation.
- Ships embodiment-specific checkpoints for YAM Box, UR AI Trainer, and FR3 Duo dual-arm robots.
- Built on a Phi-derived VLM backbone with a 12-block flow-matching action expert using grouped-query attention.
- Embodiment midtraining separates "learn the robot" from "learn the task," cutting finetuning data needs.
- Internal latent policy enables online adaptation from as few as 15 corrected episodes via FlowDAgger.
- Matches or beats existing open-weights VLAs across BusyBox, toolbox, and test-tube benchmarks.
Microsoft releases Rho, a 5B vision-language-action model for dual-arm robots
Microsoft Research has released Rho paper, an open-weights family of 5-billion-parameter vision-language-action models for bimanual robotic manipulation. A VLA combines camera input, natural-language instructions and robot state to generate physical actions. Rho separates adaptation to a robot’s hardware from training for a specific task, reducing the task demonstrations required for each deployment.
The release includes three checkpoints specialized for dual-arm platforms, along with a lightweight module that can learn from human corrections after deployment. Rho also provides the backbone for Microsoft’s previously announced Rho-alpha research model. The model weights are available through the Hugging Face collection.
| Release component | Purpose |
|---|---|
| 5B generalist VLA | Processes images, instructions and robot state to generate continuous actions. |
| Three embodiment checkpoints | Specialize the base model for the YAM Box, Franka FR3 Duo and Universal Robots AI Trainer. |
| Latent adaptation module | Updates a small internal policy from corrective demonstrations while keeping the main action generator frozen. |
Robot adaptation gets its own training stage
Task-specific VLA training commonly requires teleoperated demonstrations that teach two things simultaneously: how to control the robot and how to complete the task. Changes to the arms, cameras, grippers or control interface can force teams to collect another substantial dataset even when the target task remains similar.
Rho divides that process into distinct stages:
- Start with a visual-language backbone. Microsoft distills the model from its Phi-family vision-language models.
- Pretrain for physical interaction. The training mixture combines multi-task, multi-embodiment robot demonstrations with robotics-related Web visual question answering data.
- Specialize for one embodiment. An embodiment midtraining stage adapts observations and controls to a particular robot configuration.
- Fine-tune for the target task. Operators train the specialized checkpoint on demonstrations of the behavior they need.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.