Sakana AI Triples Robot Success Rates With SAIL
Sakana AI and the University of Tokyo have developed SAIL, a method that uses test-time search to dramatically improve how vision-language models control physical robots.

Researchers from Sakana AI and the University of Tokyo have introduced Scaling In-Context Imitation Learning, or SAIL. This new framework, accepted to the IROS 2026 conference, allows a frozen vision-language model to generate more reliable robotic movements by searching, simulating, and refining trajectory paths before executing them. By utilizing Gemini Robotics-ER 1.5 as both the trajectory generator and the evaluator, the system requires no weight updates or fine-tuning to boost performance.
In simulated testing across six ALOHA manipulation tasks—including picking up a banana, uncapping a pen, moving a bowl, opening a drawer, closing a laptop, and grasping a marker—SAIL demonstrated massive improvements. The average success rate rose from 25 percent with a single candidate trajectory to 73 percent when given a budget of 45 candidates. At a 15-node budget, SAIL achieved a 65 percent success rate, outperforming breadth-first search at 51 percent and depth-first search at 37 percent. Individually, the bowl task hit 100 percent success with six nodes, while the banana task reached 80 percent.
The researchers also validated the system on physical hardware using a LeRobot SO-101 arm for a block-in-bowl task. Operating with a 15-candidate budget, the robot succeeded in five out of six trials. Additionally, the team successfully trained a separate imitation policy on the successful trajectories generated during the search process. This offline policy also achieved a five-out-of-six success rate on the physical hardware, offering a way to bypass heavy real-time computation during deployment.
For robotics practitioners, SAIL provides a blueprint for using test-time compute to overcome the errors of one-shot planners without expensive retraining. However, the approach requires a robust simulator, a demonstration library, and a scene reconstruction pipeline. Because the final trajectory runs open-loop, the robot cannot make real-time visual corrections during movement, making the system less suitable for dynamic environments or moving objects.
This is our own summary of reporting by AlphaSignal



