Papers for

ai system safety teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Imperfect reward checks can lead to errors and fixes in reinforcement learning

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

Abstract: In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.

Mon 28 SeptArtificial Intelligence
The gist
Sometimes computers learn to do tasks by getting rewards, but the system that checks if they did a good job can make mistakes. This can trick the computer into getting rewards even when it’s wrong. The authors studied when this happens and found that it’s usually hard to catch these errors with just the usual feedback. They propose adding extra checks called audits that help the computer learn to avoid mistakes while still doing the right tasks.
Open → 2609.35677v1