Object centric learning improves robot tool use from human videos

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

RoboticsArtificial Intelligence

Summary

Robots struggle to learn complicated tool use because they lack enough real-world practice data. The authors created a method called P2P-T that teaches robots how to use tools by watching humans, without needing perfectly matching robot videos. Their system first learns stable object positions and then uses this knowledge to control the robot precisely. This approach lets robots learn faster and perform better at tricky tool tasks than previous methods.

What this means in practice

  • For robotics engineers: Teach robots precise tool handling skills by using human demonstration videos without needing paired robot data.
  • For industrial automation teams: Develop automated assembly processes where robots learn complex tool tasks efficiently from human operator recordings.

Authors

Bangjun Wang, Longyan Wu, Yukun Wei, Shenghe Shao, Chaoyi Huang, Wenze Cui, Zetong Xu, Hanlin Wu, Long Chen, Yi Ma, Hongyang Li

Abstract

Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.