What do Reward Models Memorize?

2026-07-27Machine Learning

Machine LearningComputation and Language
AI summary

The authors studied how reward models (RMs), which learn from human preferences, tend to remember certain patterns. They found that these models often focus on easy examples and memorize shortcuts that are specific to the training data, like who gave the preference or how it was collected. The models also rely too much on simple clues like response length when faced with new preference examples. This suggests that current RMs are biased and struggle to judge responses accurately in different contexts.

Reward ModelsHuman Preference DataDiscriminative TrainingMemorizationCounterfactual MemorizationBias in AIGeneralizationShortcut LearningHeuristicsContext-Dependent Evaluation
Authors
Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova
Abstract
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.