Personalized language model generation improved with efficient ranking models

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

Machine LearningComputation and Language

Summary

Personalizing language models to match different user preferences is often difficult because current methods treat all users the same. The authors found that generating many possible responses and picking the best one after seeing them can greatly improve personalization. However, scoring many responses with large models is slow and costly. To fix this, they created a small ranking model that reuses parts of the main language model to quickly score many options based on fine-tuned personal preferences. This method works better and faster than previous large models on various tasks involving personalized text generation.

What this means in practice

  • For chatbot developers: Improve chatbot responses by efficiently selecting personalized outputs from many options without heavy computation costs.
  • For recommendation system engineers: Use lightweight ranking models to quickly score many personalized content options, enabling faster and more accurate recommendations.

Authors

Qiyao Ma, Junshan Zhang, Zhe Zhao

Abstract

Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.