Agentic reward system improves robot task progress estimation

ARS: Agentic Reward System for Robot Learning

RoboticsMachine Learning

Summary

Robots need to understand how well they are doing on a task by recognizing meaningful changes in what they do. The authors created a system called the Agentic Reward System (ARS) that uses general vision-language AI models without extra training to judge task progress from videos and instructions. ARS breaks down actions into events, checks them carefully, and estimates progress frame by frame. This helps robots learn better from their experiences, especially when actions are mixed in quality or involve complex tasks like assembling machines.

What this means in practice

  • For robotics engineers: Use ARS to improve how robots estimate task progress and learn behaviors from offline video data with instructions.
  • For industrial automation teams: Integrate ARS to enhance learning of complex assembly tasks like multi-screw fastening in automated production lines.

Authors

Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng, Weiwei Zhang, Kai Xie, Xiaofeng Mou, Yi Xu

Abstract

Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars