Agents reuse historical action credit to reduce tool interactions

When Does Action Credit Need Updating?

Artificial Intelligence

Summary

For agents that use tools, updating decisions after each change can be slow and costly if they recalculate old information every time. The authors found that agents often do not need to redo all their previous assessments because small updates usually do not change which option is best. They created a way to measure when old information stays useful and a method to update it efficiently. Their system cuts down the number of times agents have to redo tool actions by nearly 40% while keeping decision quality nearly the same.

What this means in practice

  • For robotic software engineers: Reduce the number of tools or sensor calls when updating robot decision policies by reusing or correcting past action credit instead of recomputing from scratch.
  • For automated customer support developers: Minimize fresh API calls for chatbot actions after policy tweaks by updating historical action values selectively, speeding up model refresh cycles.

Authors

Hongye Yang, Boxiao Huang

Abstract

Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.