Personalized reward models improve with preference-aligned fast weights

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Artificial Intelligence

Summary

Most systems that teach computers how to prefer some answers over others use a one-size-fits-all model that ignores individual likes and dislikes. The authors found that simply giving past examples of preferences doesn’t help the model understand how those preferences relate to one another. They created a new way to quickly adjust the model’s settings during use, so it better reflects each user's unique preferences. This new method works faster and more accurately without needing extra training steps when making predictions.

What this means in practice

  • For ai product teams: Improve personalization in AI products by adapting reward models to individual users’ preference patterns during inference without extra training cost.$Commercial implications: Enables AI products with personalized feedback models that adapt quickly to individual preferences, enhancing user experience and engagement.
  • For recommendation system engineers: Use preference-aligned fast weights to better capture user-specific preference relations in recommendation models at test time for improved relevance.

Authors

Bohao Wang, Xiaoyan Zhao, Yang Zhang, Jinghang Guo, Chun Chen, Can Wang, Jiawei Chen

Abstract

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.