Sakana AI's SAIL Triples Robot Success Rates by Searching Before Moving
Sakana AI and the University of Tokyo show that a vision language model can plan robot motions more reliably by testing and refining candidate trajectories in simulation before touching the real robot.
- Sakana AI and the University of Tokyo introduce SAIL, a test-time scaling method for VLM robot control, accepted to IROS 2026.
- MCTS over full trajectories, with each node a candidate motion and each edge a refinement guided by VLM feedback.
- Average success across six ALOHA manipulation tasks climbs from 25% at one candidate to 73% at 45 candidates.
- Uses Gemini Robotics-ER 1.5 as both policy and evaluator, with no weight updates at any stage.
- Physical LeRobot SO-101 block-in-bowl task succeeded in five of six trials using a 15-candidate budget.
- Open-loop execution and heavy simulation compute remain the main limitations. Paper.
SAIL searches its way to better robot plans
Researchers at Sakana AI and the University of Tokyo report that a frozen vision-language model can produce more reliable robot motions by searching, simulating, and revising candidate trajectories at inference time. Their method, SAIL, raises the average success rate across six simulated manipulation tasks from 25% with one candidate to 73% with 45. The paper has been accepted for IROS 2026.
SAIL, short for Scaling In-Context Imitation Learning, offers developers a way to improve an unreliable one-shot planner when they have a simulator, a demonstration archive, and enough latency budget for search. The model weights remain fixed throughout the process.
Why a single trajectory breaks
In-context imitation learning gives a robot examples of successful behavior, then asks a model to solve a new scene from those demonstrations. A vision-language model can inspect the scene and generate a complete sequence of robot-hand positions, orientations, and gripper commands. Small changes in object placement, sampled outputs, or retrieved examples can still produce large differences in execution.
A grasp that misses by a few millimeters can invalidate every later waypoint. SAIL applies test-time scaling, the practice of spending more computation on each input while keeping model parameters fixed, to physical trajectories. It generates several plans, evaluates their simulated execution, and uses the results to revise promising candidates.
Inside SAIL’s search loop
SAIL places Gemini Robotics-ER 1.5 inside a Monte Carlo tree search. The model serves as both trajectory generator and evaluator. Each tree node contains a complete trajectory, while each edge represents a refinement. The search can explore new revisions and continue developing branches that receive higher scores.
- Retrieve and generate. The system queries an archive of successful trajectories for scenes resembling the current object arrangement. The model receives those demonstrations, the current scene, and the robot state, then generates a full candidate trajectory.
- Simulate and score. A simulator executes the candidate and records a video. The evaluator estimates progress through predefined subtasks and converts that progress into a scalar score. Partial progress can therefore distinguish two failed candidates. The scoring design draws on ROVER, which tracks task completion over time.
- Align feedback with waypoints. Scores from sampled video frames are mapped back to specific trajectory steps. The generator receives that step-level feedback and revises sections where progress stalled while preserving successful motion segments.
- Select for execution. Candidate testing remains in simulation. The physical robot receives the selected trajectory after the search finishes.
Implementers must provide the demonstration archive, simulator, scene reconstruction pipeline, and subtask rubric used by the evaluator. SAIL supplies the search and revision structure around those components.
More nodes produce more successful plans
The researchers evaluated six tasks in the ALOHA simulator, using 20 initial configurations for each task. The benchmark covered picking up a banana, uncapping a pen, moving a bowl, opening a drawer, closing a laptop, and grasping a marker.
| Method | Search nodes | Average success |
|---|---|---|
| Single generation | 1 | 25% |
| Breadth-first search | 15 | 51% |
| Depth-first search | 15 | 37% |
| SAIL | 6 | 55% |
| SAIL | 15 | 65% |
| SAIL | 30 | 71% |
| SAIL | 45 | 73% |
At a budget of 15 nodes, SAIL reached 65% average success, compared with 51% for breadth-first search and 37% for depth-first search. The bowl task reached 100% success with six nodes, while the banana task reached 80%. These matched-budget results indicate that retrieval and waypoint-level feedback contribute beyond the number of sampled plans.
One real arm, six trials
For the hardware test, the team used a LeRobot SO-101 arm for a block-in-bowl task. Color and depth cameras supplied the data needed to reconstruct the scene in simulation. SAIL searched 15 candidate trajectories, selected one, and sent it to the arm.
The search pipeline succeeded in five of six trials. The authors attribute the remaining failure to pose-estimation errors and differences between simulated and physical contact dynamics.
A separate imitation policy trained on successful trajectories generated by the search also completed five of six trials. That approach moves the computational cost into data generation and training, allowing deployment without running tree search for every attempt.
Whole trajectories become search objects
SAIL’s central contribution is trajectory-level test-time scaling. The search operates on complete robot motions, and the evaluator’s temporal scores identify which sections need revision. This structure lets a frozen foundation model use additional inference compute to handle object arrangements that defeat its first prediction.
The paper’s ablation results support the value of both retrieval and detailed feedback. Random retrieval and weaker feedback plateau as the node budget grows, while the complete pipeline continues to improve through 45 nodes.
Where the method fits
- Suitable workloads: Tabletop manipulation tasks with a useful demonstration library, a simulator that approximates scene geometry and contact, and enough time to search before execution.
- Control constraint: The selected trajectory runs open-loop. The robot receives no visual correction while moving, which limits the approach on moving objects, uncertain contacts, and tasks that require continuous replanning.
- Infrastructure cost: Each candidate requires simulation, video evaluation, and potentially another model call for refinement. Larger search budgets improve success while increasing compute and latency.
- Transfer risk: Errors in pose estimation, geometry, or contact modeling can cause a plan that succeeds in simulation to fail on hardware.
- Evidence limit: The simulated evaluation covers six tasks, while the physical validation covers one task with six trials for each tested method.
The paper and project page include simulated and physical rollouts. For robotics teams already using vision-language models, SAIL provides a concrete architecture for adding retrieval, simulation-based scoring, and iterative trajectory repair without fine-tuning the underlying model.