Memory augmented AI struggles to apply user preferences correctly
PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use
Artificial Intelligence
Summary
AI helpers sometimes remember user preferences but don’t always know when to use them. The authors created a test called PairPref that checks if AI can tell when a preference should affect the response based on the situation. They found many AI models often apply preferences even when it’s not appropriate. This shows AI still has trouble judging the right context to use what it remembers.
What this means in practice
- •For assistant developers: Improve AI assistants by evaluating when to apply user preferences based on context using the PairPref benchmark.
- •For chatbot trainers: Train chatbots to better decide when to use stored preferences in conversation with the new PairPref test.
Authors
Mingfei Lu, Mengjia Wu, Yi Zhang
Abstract
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less attention to when those preferences should apply. We introduce PairPref, a benchmark of contextual preference use. Each pair changes only the situation, keeping the preference, request, and four candidate replies fixed. The preference remains valid in both situations. In the selection track, models must choose the reply that applies the preference only where appropriate. In the free-generation track, they must decide when to apply it without seeing candidate replies. Both tracks use the same 1,227 pairs across 45 preferences and eight situation categories. We evaluate eight models, most of which achieve selection scores ($Δ$) of 51 to 65 points. In free generation, however, both responses are appropriate for their respective situations in only 3.6\% to 18.3\% of pairs. Models continue to apply the preference in both situations even with fewer retrieved memories, alternative presentation formats, and a stricter prompt. These results show that models still struggle to judge when user preferences apply and respond accordingly.