Llms often apply preferences when they should not during personalized responses

Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs

Computation and Language

Summary

Personalized language models try to decide when to use stored user preferences for each generated response. The authors found that these models often mistakenly apply preferences even in contexts where they should not, leading to over-personalization. They broke down the problem into knowing if a preference is relevant, deciding explicitly whether to apply it, and then generating a response accordingly. Their analysis showed the main issue arises in the decision step, not in understanding or generating language. They developed a method to detect a bias toward applying preferences and showed that adjusting for this bias reduces errors while keeping responses accurate.

What this means in practice

  • For ai system developers: Improve personalized language models by correcting decision biases to reduce incorrect application of user preferences during response generation.
  • For customer support teams: Enhance automated support chatbots to better follow user preferences only when relevant, reducing customer confusion from over-personalized replies.

Authors

Haeun Jang, Yonghyun Jun, Hwanhee Lee

Abstract

Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot tell where this failure arises. We decompose preference handling into three stages and measure each separately: (1) knowing whether a preference applies, (2) deciding on an explicit Apply/Suppress label, and (3) generating a response consistent with that label. Using linear probes, we first show that this applicability signal remains decodable from hidden states during generation. By making the decision explicit, we then find that in most settings wrong decisions faithfully followed outnumber correct decisions lost in generation. We thus locate the failure in the decision, which breaks once the model is also asked to answer. To determine whether this reflects lost sensitivity or a response bias, we propose ABIDE (Apply-Bias Investigation via Decision-score), which adapts signal detection theory to Apply-vs-Suppress decision scores read directly from logits. ABIDE reveals a generation-induced Apply bias: merely stating an answer-generation objective shifts the decision score toward Apply while sensitivity is largely preserved, and the shift persists under controls for prompt structure, cascades across preference slots, and prompt wording. Finally, we show that subtracting a single bias scalar, estimated on a held-out split, from the decision score at decoding time reduces leakage while largely preserving fulfillment.