Mobile robots learn complex two-hand tasks from human whole-body videos

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

Robotics

Summary

Teaching robots to use both hands while moving around is very complicated because it needs smooth coordination of walking, whole-body movement, and finger movements at the same time. The authors created DexRoam, a system that learns these skills by watching humans perform tasks while wearing simple VR gear, without needing extra cameras or trackers. They carefully align the human movements to robot actions to keep important details intact. Their tests show that adding human demonstrations significantly improves robot learning, cutting down the number of robot trials needed and boosting success rates.

What this means in practice

  • For robotics engineers: Train mobile robots to perform fine and coordinated two-handed tasks by using human whole-body motion captured with consumer VR equipment.
  • For industrial automation teams: Reduce physical robot training trials in factories by leveraging scalable human demonstration data for complex manipulation tasks.

Authors

Rui Zhou, Yibo Yuan, Junkai Zhao, Fangyuan Zhao, Xiaoguang Zhao, Shanghang Zhang, Sirui Han

Abstract

Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simplifying human motion, discarding exactly the fine-grained, coupled structure such tasks depend on. We present DexRoam, a complete system for learning mobile bimanual dexterous manipulation from human demonstrations, in which whole-body motion remains continuous and coupled throughout the human-to-robot transfer process. To enable scalable collection of whole-body human manipulation demonstrations, we develop a tracker-free capture system using only a consumer VR headset and a head-mounted stereo camera, without external cameras or motion trackers. We then perform three explicit alignment stages---embodiment, action-semantic, and temporal---to map captured motion into the robot action space, preserving fine-grained whole-body motion and allowing human and robot demonstrations to be jointly learned by standard VLA policies. Real-world experiments with different VLA backbones show that human demonstrations consistently improve policy learning across training paradigms, raising average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5, while matching robot-only training with half the robot demonstrations. Ablations confirm that each alignment stage is necessary. These results highlight the potential of human demonstrations for scalable whole-body mobile manipulation with preserved fine-grained motion structure.