How reward objectives and target assignments shape learned preferences
From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
Artificial IntelligenceMachine Learning
Summary
Training AI systems often involves showing them preferences with different strengths and then turning those into rewards to guide learning. This paper studies how the way these preference strengths are assigned affects the rewards produced by various methods. The authors found that keeping the assignment correspondence intact maintains clearer preference differences under equal accuracy. Different reward methods also react differently to the same preference assignments. They offer new ways to evaluate how target placement and reward choices interact beyond just accuracy.
What this means in practice
- •For machine learning engineers: Design reward functions in preference-based models to improve preference margin clarity using new assignment strategies.
- •For natural language processing developers: Improve fine-tuning of language models by selecting reward objectives and target assignments that optimize reward signal properties beyond accuracy.
Authors
Jiangtao Lin, Bangyang Wei, Siyi Liu, Yihang Ding, Yuhan Dong
Abstract
Soft preference targets specify supervision strength, and reward objectives convert that strength into learned reward signals. A central design question remains: how does assigning a fixed set of preference strengths to different response pairs change the rewards produced by different objectives? We introduce assignment geometry to study this interaction. Mean-matched smoothing controls target dispersion, while within-stratum reassignment changes correspondence and preserves the complete target distribution. Across five reward objectives, intact correspondence retains the largest clean preference margins among the compared soft targets within a common accuracy-equivalence budget. Attenuation orderings change with the reward objective, revealing different responses to the same target assignments. Independent reassignments and a related source construction reproduce the retention direction. An attenuation-retention profile compares these combinations through margin magnitude, edit response, and accuracy. Against independently calibrated scaling, APLOT uniform targets deliver additional attenuation on both aggregate and presentation edits. These findings establish a joint design space in which target placement and reward objective shape reward properties beyond preference accuracy.