Attentive search improves credit assignment in agent workflows

ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning

Artificial Intelligence

Summary

Making decisions in long workflows is hard to learn from because it’s unclear which actions caused success. The authors propose a method called ASCT that looks at alternative actions at each decision point using a special tree search. This helps assign credit to better actions while still using the original action choices to train. They tested ASCT in a question-answering system and showed better performance and efficiency than past methods. This approach helps learn more precise policies for complex workflows.

What this means in practice

  • For ai system engineers: Improve training of decision-making agents in multi-step tasks by better assigning credit to individual actions using counterfactual tree search.
  • For machine learning platform teams: Reduce resource use and improve efficiency in policy learning systems by integrating attentive counterfactual evaluation at training time.

Authors

Yang Li, Jinhan Yang, hai liu, Di Wan, Xiyu Chen, Zongsi Xu, Tuo Zhou, Sheng Zhong, Sergey Volkov, Ye Luo, Hao Sun

Abstract

Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tree evaluates alternative legal actions from the same recoverable prefix. Its action-value table is centered by the frozen actor's probabilities and supplies credit for PPO on actor-sampled trajectories. This protocol connects counterfactual evaluation to policy learning while deploying the actor alone. Uniform, UCT, and cost-aware AgentUCT instantiate the framework. On HotpotQA agentic retrieval-augmented generation, all three improve mean held-out utility over trajectory-return PPO and workflow-adapted VinePPO. Across three seeds, ASCT-AgentUCT reaches 0.6187 utility versus 0.5939 for VinePPO, with gains in answer F1 and execution cost, and uses 50.3% fewer recorded auxiliary Qwen tokens. Transfer and component-description studies examine the learned policies beyond the training setting.