Humanoid robots learn to carry diverse objects from few videos

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

RoboticsComputer Vision and Pattern RecognitionGraphics

Summary

Teaching humanoid robots to carry and manipulate objects by watching videos is hard because it's difficult to get many good examples. The authors developed a method called PRISM that takes just a few real videos and creates many new, varied ones by changing object interactions. These videos are then turned into realistic motions, allowing the robot to learn how to carry different objects it hasn't seen before. The robot successfully picks up and carries items like boxes and balls without extra real-world training.

What this means in practice

  • For robotics engineers: Train humanoid robots to handle new objects by expanding limited video data into diverse scenarios for better manipulation skills.
  • For industrial automation teams: Deploy humanoid robots that can pick up and carry varied items in warehouses without requiring extensive new training data for each object.$Commercial implications: Enables sale of adaptable humanoid robots suitable for logistics by reducing costly data collection and fine-tuning.

Authors

Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan, Jitendra Malik, Carmelo Sferrazza, C. Karen Liu, Guanya Shi, Angjoo Kanazawa

Abstract

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.