Imperfect reward checks can lead to errors and fixes in reinforcement learning
Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
Artificial Intelligence
Summary
Sometimes computers learn to do tasks by getting rewards, but the system that checks if they did a good job can make mistakes. This can trick the computer into getting rewards even when it’s wrong. The authors studied when this happens and found that it’s usually hard to catch these errors with just the usual feedback. They propose adding extra checks called audits that help the computer learn to avoid mistakes while still doing the right tasks.
What this means in practice
- •For machine learning engineers: Improve training of AI systems by integrating audits to reduce errors caused by imperfect reward feedback.
- •For ai system safety teams: Develop better methods to detect and fix reward hacking in reinforcement learning setups.
Authors
Christian Moya, Elliott Thornley, Guang Lin
Abstract
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.