Language models adopt false claims differently based on source authority

Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation

Computation and Language

Summary

Large language models can wrongly accept false information from earlier parts of a conversation, which the authors call session-level contamination. They tested this by putting false statements into different parts of a chat and seeing if models treated them as true. Some models, like GPT-5.4 Mini, never accepted these false ideas, while others like Gemini and GLM models accepted them more depending on who the false information was attributed to. This shows that how much a model trusts a source affects its mistakes and recovery from them. The authors suggest systems should track where information comes from to avoid being tricked.

What this means in practice

  • For ai platform engineers: Design language model systems that track information sources to prevent adoption of false context in multi-turn conversations.
  • For customer support developers: Improve chatbot reliability by understanding how different models respond to false information injected in conversation turns.

Authors

Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah

Abstract

Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct failure mechanisms while holding the false premise constant, and evaluate GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains at temperature zero (22,500 turns), using a dual-track automated judge validated against a human gold standard (Cohen's \k{appa} = 0.901). GPT-5.4 Mini showed zero adoptions across all 500 sessions, a content-independent policy at the session level; token-level probing shows the underlying margin, while large, is finite. Gemini-3.1 Flash-Lite followed a steep authority gradient: 0.1% adoption for self-attributed falsehoods, 23.5% for user-cited sources, 68.2% for system-injected authority, and 94.0% under instruction override. GLM-4.5-Air showed a shallower gradient (15.8% vs 84.2%), a 68-percentage-point dissociation confirming that authority deference and instruction compliance are distinct mechanisms within one architecture. Recovery also diverged: GLM recovered in 94.5% of affected sessions, whereas 26.1% of affected Gemini sessions never did, rising to 40.0% under instruction override. Conversation history is an untrusted attack surface requiring provenance-aware system design; the complete framework is released as an open-source benchmark.