Robots learn reusable skills better using policy structure cues
Agent Priors-guided Policy Learning
Robotics
Summary
Robots often learn tasks by combining smaller skills, but these skills sometimes don’t work well in new situations or when combined differently. The authors show that telling the robot about how each skill is supposed to work—like knowing a grasp only depends on the gripper’s position relative to an object—helps it learn and choose the right skills better. They built a system called APPL that breaks down demonstrations into skills, tries different ways to explain each skill, and then picks the best approach when doing new tasks. This method improved robot performance on tests that required using skills in new ways.
What this means in practice
- •For robotics engineers: Enable robots to reuse trained skills more reliably in new combinations and unfamiliar scenarios during task execution.
- •For automation system integrators: Build flexible automated systems that adapt by selecting context-aware robot behaviors from multiple learned policies.
Authors
Puming Jiang, Tianrun Hu, Haozhe Du, Yibo Li, Zhiwei Xue, Xinhu Li, Harold Soh
Abstract
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.