DexAgent generates adaptable robot skills from human videos for dexterous tasks
DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library
Robotics
Summary
Teaching robots to do tricky tasks with their hands is hard because objects can move or change shape. The authors created DexAgent, a system that watches a single person’s video doing a task and turns it into robot movements. It checks its work at every step to avoid mistakes and can create new robot skills when needed. Over time, DexAgent learns and gets faster by remembering what it discovered. Robots trained this way succeeded much more often on real tasks than before.
What this means in practice
- •For industrial robot integrators: Create tailored robot manipulation skills by converting single videos of humans manipulating diverse parts into robot actions adaptable to different objects.
- •For virtual reality developers: Generate realistic robot hand interactions from human action videos to enhance simulation environments with physically valid object manipulations.
Authors
Youhui Wang, Yunzhu Li, Li Fei-Fei, Jiajun Wu, Huang Huang
Abstract
Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformable objects. We introduce DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. It operates through four stages: semantic understanding of human videos, property-based simulation reconstruction, robot trajectory optimization, and robot data generation. At each stage, DexAgent adapts its approach to the task and object properties by selecting suitable skills from its tool library or developing new ones when needed. Property-specific verifiers assess stage outcomes for physical validity and task-specific requirements and provide feedback for refinement, preventing error propagation through the workflow. This adaptive, verification-guided process allows DexAgent to process diverse objects and long-horizon tasks. In the final stage, DexAgent varies object and robot states in simulation to generate diverse robot trajectories from a single human video, then retextures the rendered observations to facilitate sim-to-real transfer. Newly developed skills and verifiers are retained in its tool library, making it self-evolving to accumulate reusable capabilities. This reduces processing time as DexAgent encounters more human videos. Across eleven real-world tasks, policies trained with DexAgent-generated data achieve a 3.5x higher success rate than competing baselines. Project website: https://dexagent123.github.io/.