GraphHCA improves credit assignment for long-horizon AI tasks

GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents

Machine Learning

Summary

When teaching AI agents that use large language models to achieve long goals, it’s hard to tell which individual actions helped reach the final success. The paper introduces GraphHCA, a new math-based method that assigns credit to each step without extra complex models. This approach uses probabilities from past attempts to better understand which moves made a difference and improves AI performance on several tasks. The authors show it works better than previous methods on some challenging environments.

What this means in practice

  • For ai engineers: Improve step-level feedback in training language-based agents for complex, multi-step tasks where success signals are sparse.
  • For robotics teams: Enhance control policies in robots performing sequential goal-directed tasks using vision and language inputs by better credit assignment.

Authors

Haodong Zhu, Yangyang Ren, Changbai Li, Sheng Xu, Linlin Yang, haiguang liu, Baochang Zhang

Abstract

Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not explicitly capture its retrospective relation to the realized outcome. Hindsight credit assignment (HCA) instead attributes credit through the ratio of hindsight to behavior-policy probabilities, but estimating the hindsight distribution requires an auxiliary model or an extra pass. To address this estimation bottleneck, we propose GraphHCA, a model-free realization of HCA that eliminates explicit hindsight-distribution estimation. For terminal-goal tasks with deterministic transitions, Bayes' rule reduces the hindsight ratio to a ratio of behavior-policy success probabilities at consecutive states. Taking logs yields a state-wise success potential, whose increment across a transition provides step-level credit. GraphHCA estimates this potential from pooled rollouts through a discounted recursion on the induced transition graph, which admits a unique fixed point on any directed graph. The resulting step-level signal is combined with the trajectory-level advantage, requiring neither a learned hindsight model nor an extra forward pass and recovering GRPO when the step-level weight is zero. Among all compared baselines, GraphHCA achieves state-of-the-art results on ALFWorld and WebShop at both LLM scales, and on Sokoban with a vision-language agent. For example, on ALFWorld it improves overall success rate by up to 24.6 points over GRPO and by up to 4.7 points over the strongest step-level baseline.