X-Reset trains robot hands to grasp diverse objects using human resets
X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
Machine LearningArtificial IntelligenceRobotics
Summary
Training robot hands to pick up and use many different objects is hard because robots struggle to learn these skills from scratch. The authors propose X-Reset, a method that uses snapshots from human hand motions interacting with objects to help robots start their learning in helpful positions. Instead of copying human movements exactly, X-Reset converts these human poses into robot poses and carefully chooses stable states to reset to during training. This approach allows robots to learn general skills for many objects across different types of robot hands, even working with imperfect human data and transferring skills from simulation to real robots.
What this means in practice
- •For robotic automation teams: Train robot arms with different grippers to handle a wide variety of objects without task-specific programming.
- •For industrial robot integrators: Use human demonstration data to speed up robot learning and deployment in new manipulation tasks across product lines.$Commercial implications: Enables sale of adaptable robot learning solutions that reduce demonstration costs and improve deployment speed.
Authors
Prithwish Dan, Chenyang Ma, Wei Zhan
Abstract
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.