Agentic reinforcement learning improves with unified feedback arbitration

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Artificial Intelligence

Summary

Reinforcement learning agents often struggle to learn well when rewards are sparse and delayed because it's hard to figure out which actions led to success. The authors found that two types of feedback—outcome-based and hindsight-based—sometimes disagree on which decisions were good. They created a method called UniOPSD that smartly combines these feedback types by deciding how much to trust each one at every step. This approach improved the agents’ performance on various challenging tasks.

What this means in practice

  • For autonomous system developers: Create AI agents that learn better by adaptively combining different types of feedback for more accurate decision credit assignment.
  • For robotics engineers: Improve robot policies by integrating outcome and hindsight information to accelerate learning in complex, long-horizon tasks.

Authors

Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Xiaofeng Han, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, Minghui Wu, Guannan He, Changwei Wang

Abstract

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of $82.8\%$ and $83.6\%$, WebShop success rates of $75.0\%$ and $82.0\%$, and Search-QA aggregate accuracies of $45.3\%$ and $49.8\%$, respectively. On 3B WebShop, UniOPSD improves over SDAR by $7.0$ percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd