Learning When to Trust via Selective Context Preference Optimization

2026-08-06Computation and Language

Computation and LanguageArtificial IntelligenceMachine Learning
AI summary

The authors explain that language models often rely on extra information to answer questions, but sometimes a single wrong piece can make them give a wrong answer. Instead of ignoring all extra info to avoid mistakes—which makes the model useless when the info is actually helpful—they suggest the model should learn when to trust or ignore it (called selective trust). They created a test called MIST to measure this and introduced a method named SCOPE to train models to do better at it. Their approach helps models avoid being tricked by misleading info while still using good information correctly.

language modelscontext conditioningselective trustmisleading signalsbenchmarkDirect Preference Optimization (DPO)model robustnesspreference learningopen-sourced modelspaired evaluation metrics
Authors
Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
Abstract
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open-sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.