Robot skill learning from human videos measured with a new benchmark
Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation
Robotics
Summary
Teaching robots to learn tasks just by watching humans is tricky because current tests and setups vary a lot, making it hard to compare progress. To fix this, the authors created RoboReel, a set of videos, robot actions, and tasks that lets different robot learning methods be tested fairly on ten different manipulation jobs. They tested several types of learning methods and showed that robots still struggle with longer tasks and ones requiring very precise actions. This benchmark helps people see which approaches work best and where more improvements are needed.
Learning from Observationrobot manipulationbenchmarkhuman demonstrationsimulationlong-horizon tasksvisual distractorsrobot policy learning
Authors
Weiwei Gu, Anmol Gupta, Anant Sah, Ryan Varghese, Lalitha Shreya Vanam, Prabhath Adireddi, Peter Karkus, Nakul Gopalan
Abstract
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models' performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models. Webpage: https://roboreel.github.io