Robot skills improve by practicing tasks found in past data
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
RoboticsArtificial Intelligence
Summary
Building robots that can reliably do many different tasks usually requires a lot of human work to teach and control them. The authors created a method where a robot can improve its skills on its own by practicing tasks it identifies from previous experiences in simulation. Their method helps the robot learn new helpful skills and fix mistakes without changing its main program. This process makes robots much better at completing tasks, even in the real world, after practicing many times in simulation.
What this means in practice
- •For robotics engineers: Improve robot task reliability by enabling autonomous skill refinement using past task data without retraining core models.
- •For automation system developers: Create automation systems that adapt to hardware changes through simulated practice and execution feedback before real deployment.
Authors
Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi
Abstract
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/