Action-conditioned tactile models improve robot hand manipulation learning

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

RoboticsComputer Vision and Pattern Recognition

Summary

DexTouch-WM helps robots learn how to handle objects by predicting what will happen when they touch and move things. Instead of only using data from robots, the authors collect touch data from human hands using similar sensors and make that information match the robot's movements. This shared data helps the robot build better predictions about how objects and touch sensors will behave, even when humans and robots try different tasks. The approach improves robot learning and can generate simulated experiences for further training.

What this means in practice

  • For robot developers: Train robot hands to predict touch and visual outcomes using combined human and robot tactile data to improve manipulation skills.
  • For robotics simulation engineers: Use learned tactile world models as realistic surrogate environments to evaluate and generate training data for robot manipulation policies.

Authors

Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen, Renjing Xu

Abstract

Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.