ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm
2026-07-27 • Robotics
Robotics
AI summaryⓘ
The authors created ArmnetBench v0.1, a system for testing robot arm policies using many low-cost robot setups with minimal human help. They evaluated seven different robot control methods across twelve tasks, collecting over 3,000 episodes labeled by success level. Their dataset includes human-judged policy attempts and perfect demonstration runs, useful for training and comparing future robot learning methods. The authors also released all their data publicly to encourage more research and fair comparisons.
robot manipulationbenchmarkpolicy rolloutsdemonstrationsbimanual controlhuman scoringrobot learningdata labelingevaluation metricsrobot arm
Authors
Praveen Selvaraj, Lorenzo Uttini, Ville Kuosmanen
Abstract
Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout requires physical hardware and an operator to set up, reset, and score it. We introduce ArmnetBench v0.1, a benchmark run on a fleet of low-cost SO-101 cells under light on-site supervision. v0.1 validates this arm farm end to end and compares 7 policies across 12 tasks with both single-arm and bimanual configurations. Each policy is trained or fine-tuned on 50 demonstrations per task; the benchmark contains 2,518 policy rollouts and 600 reference demonstrations. All 3,118 episodes carry a three-way label (successful, suboptimal, or failure). Policy rollouts are human-scored, while demonstrations are successful by construction. Beyond evaluation, its quality-labelled trajectories support downstream learning, from reward and predictive world models to policies trained on mixed-quality data. The leaderboard is an initial comparison under this shared budget. We release the 3,118 core episodes in LeRobot v3.0 and RoboMeter formats.