Cross-rollout method improves long-horizon learning for AI agents
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
Machine Learning
Summary
Training AI agents to make long-term decisions is tricky because it’s hard to figure out which early actions lead to later success. The authors introduce a new method called Cross-Rollout Bellman Closure that combines information from multiple trial runs to better estimate the value of each step. This approach helps the AI learn more efficiently by using shared states in different runs to improve credit assignment. Testing on various benchmarks showed consistent improvements in performance and learning speed.
What this means in practice
- •For ai developers: Improve training efficiency of reinforcement learning agents in complex long-step tasks using cross-run credit assignment.
- •For robotics teams: Enhance robot decision-making in extended action sequences by better aggregating evidence from multiple trial runs without extra environment tests.
Authors
Yangyang Ren, Haodong Zhu, Linlin Yang, Sheng Xu, Peichao Lai, Baochang Zhang
Abstract
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. We introduce Cross-Rollout Bellman Closure (CRBC), which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest evaluated baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct.