Summary
Robots need to understand both sight and touch to handle objects like people do, but videos of humans only show what to see, not what to feel. The researchers developed a system called DEX-X that fills in the missing touch information by recreating the hand and object movements in a computer simulation. This simulation adds realistic touch feedback, which helps teach robots how to manipulate things using both vision and touch. The trained robots can then perform a variety of hand tasks in the real world without extra training, even with new objects they have never seen before. This approach suggests we can use many human videos on the internet to teach robots complex hand skills more easily.
dexterous manipulationvisual-tactile learningsimulationrobot hand-arm platformzero-shot learningsim-to-real transfermonocular human videotactile sensingpoint-cloud observationscontact dynamics
Authors
Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang, Botian Xu, Shiqin Tong, Jiajun Liu, Mingzhi Pei, Chenyu Zhang, Wanli Xing, Kaifeng Zhang, Mengdi Xu
Abstract
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.