Papers for

robotic software engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Agents reuse historical action credit to reduce tool interactions

When Does Action Credit Need Updating?

Abstract: Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.

Thu 24 SeptArtificial Intelligence
The gist
For agents that use tools, updating decisions after each change can be slow and costly if they recalculate old information every time. The authors found that agents often do not need to redo all their previous assessments because small updates usually do not change which option is best. They created a way to measure when old information stays useful and a method to update it efficiently. Their system cuts down the number of times agents have to redo tool actions by nearly 40% while keeping decision quality nearly the same.
Open → 2609.29007v1

IncentRL balances guidance and task success in reinforcement learning

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Abstract: Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performance. IncentRL adds a Kullback--Leibler (KL) penalty between a specified outcome distribution and a preferred distribution. For finite discounted Markov decision processes with bounded shaping costs, we derive an external-value perturbation bound, establish a sufficient strict-action-gap condition for preserving the original optimal policy, and characterize the large-weight regime through discounted cumulative preference cost. Exact examples clarify the limits of these guarantees, including tied optima and support mismatch. We study a practical implementation using a hand-designed, distance-based outcome proxy, a fixed preference distribution, and score-weighted coefficient search. On MiniGrid DoorKey-8x8, the reported three-seed mean success rate after two million training steps reaches 98\% with coefficient 0.01, compared with 90.5\% for the reported zero-coefficient baseline, while the search progressively shifts toward smaller coefficients. Together, these results provide a principled view of the central trade-off in preference-based RL: using additional guidance to improve learning without excessively distorting the original task objective. The current experiments remain descriptive and do not yet isolate KL shaping from simpler alternatives.

Fri 18 SeptMachine Learning
The gist
Reinforcement learning often uses extra signals to guide training, but these signals can accidentally change what the system aims to do. The authors introduce IncentRL, a method that adds guidance while measuring how much it changes the original goal. Their math shows when the original goal stays safe and how big guidance signals affect results. They test their idea on a simple game, improving success rates by carefully choosing how strong the guidance is. This work highlights how to use extra hints without messing up what the system is supposed to learn.
Open → 2609.21525v1