Summary
Getting robots to learn from videos recorded by people wearing cameras is promising, but not all video data is equally helpful. The authors studied how different types of video data—such as how well the demonstrated actions match a robot’s abilities, how long the videos are, and whether action labels are included—affect how well robots can learn new tasks. They found that videos closely aligned with robot actions make learning more effective and reduce the amount of extra robot data needed. Also, even videos without detailed action labels can still be useful for training robots. This research helps clarify what kind of human video data is most valuable to improve robot learning.
What this means in practice
- •For robotics engineers: Use aligned human egocentric videos to improve robot task learning efficiency and better generalize to new scenarios.
- •For video data annotators: Prioritize collecting and organizing video-only egocentric footage to reduce dependency on costly action labels in robot training datasets.
Authors
Zhihao Sun, Liu Liu, Xinjiang Wang, Haoyi Jiang, Wei Feng, Huiqiang Zhang, Xiaosong Jia, Zhizhong Su, Zuxuan Wu
Abstract
Egocentric human data provides a scalable source of experience for robot learning, but varies substantially in human-robot alignment, behavioral coverage, and available supervision. Existing work shows favorable scaling with increasing human data, but it remains unclear which data properties drive downstream robot gains and how to use such data throughout the training pipeline. We present a systematic study of egocentric human data with different alignment and supervision under a unified world-action model framework. With the model backbone fixed, we disentangle the effects of human-robot alignment, data duration and task diversity, action supervision, and data usage strategies. We find that aligned human demonstrations substantially improve out-of-distribution generalization and reduce target-task robot data requirements; data duration and task diversity affect downstream capabilities differently; and video-only supervision remains effective without action labels, providing a strong foundation for subsequent video-action training. We validate these findings through closed-loop policy evaluation on both real robots and RoboDojo. Rather than treating data duration as the sole scaling axis, Ego4WAM shows how alignment, task diversity, available supervision, and usage strategy jointly shape the value of egocentric human data for robot learning.