IncentRL balances guidance and task success in reinforcement learning
IncentRL: The Trade-Off Between Preference Guidance and Task Performance
Machine Learning
Summary
Reinforcement learning often uses extra signals to guide training, but these signals can accidentally change what the system aims to do. The authors introduce IncentRL, a method that adds guidance while measuring how much it changes the original goal. Their math shows when the original goal stays safe and how big guidance signals affect results. They test their idea on a simple game, improving success rates by carefully choosing how strong the guidance is. This work highlights how to use extra hints without messing up what the system is supposed to learn.
What this means in practice
- •For autonomous system developers: Improve training outcomes in reinforcement learning agents by tuning preference guidance without distorting original tasks.
- •For robotic software engineers: Use controlled reward shaping to speed up robot learning while maintaining the intended task behavior.
Authors
Xuening Wu, Yanlan Kang, Shenqin Yin
Abstract
Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performance. IncentRL adds a Kullback--Leibler (KL) penalty between a specified outcome distribution and a preferred distribution. For finite discounted Markov decision processes with bounded shaping costs, we derive an external-value perturbation bound, establish a sufficient strict-action-gap condition for preserving the original optimal policy, and characterize the large-weight regime through discounted cumulative preference cost. Exact examples clarify the limits of these guarantees, including tied optima and support mismatch. We study a practical implementation using a hand-designed, distance-based outcome proxy, a fixed preference distribution, and score-weighted coefficient search. On MiniGrid DoorKey-8x8, the reported three-seed mean success rate after two million training steps reaches 98\% with coefficient 0.01, compared with 90.5\% for the reported zero-coefficient baseline, while the search progressively shifts toward smaller coefficients. Together, these results provide a principled view of the central trade-off in preference-based RL: using additional guidance to improve learning without excessively distorting the original task objective. The current experiments remain descriptive and do not yet isolate KL shaping from simpler alternatives.