Llms vulnerable to misinformation when memory is wiped between talks
Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
Computation and Language
Summary
Large language models are used to find information and answer questions, but they can be tricked by someone intentionally giving false or misleading information. The authors found that when these models don’t remember previous conversations, they are much easier to persuade with wrong facts. They created a way to test this by wiping the model’s memory each time it talks, showing that simple tricks succeed almost all the time. Interestingly, complex attacks sometimes backfire and make the model stick to good facts instead.
What this means in practice
- •For chatbot developers: Evaluate dialogue systems for vulnerabilities to misinformation when conversation memory is disabled or reset.
- •For online platform security teams: Design more secure knowledge retrieval tools by identifying when stateless models can be easily deceived.
Authors
Zhuoang Cai
Abstract
As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a critical safety concern. Existing red-teaming frameworks typically evaluate models in multi-turn dialogues where the target model retains full conversation history. We identify a critical flaw in this setting termed \textbf{``Refusal Inertia''}: a model's initial refusal often propagates through subsequent turns largely to maintain contextual consistency, thereby masking its true vulnerability to sophisticated, isolated persuasion attempts. To rigorously evaluate the ``cold-start'' defense capabilities of SOTA models, we introduce the \textbf{SAST-IR} (Stateful Attacker, Stateless Target - Iterative Refinement) framework. By enforcing a memory wipe on the target while retaining the attacker's history, we simulate a worst-case adversarial setting using \textbf{multi-turn} (stateless) iterations. Leveraging \textbf{CP-Agent} (Cognitive Persuasion Agent), an enhanced diagnosis-guided agent, our experiments on the custom \textsc{CounterFact-Strict} dataset ($N=50$) yield alarming results: simple, diverse attack strategies achieved a staggering \textbf{96\%} success rate, exposing severe brittleness in memory-less defense. Furthermore, we reveal a \textbf{``Complexity Paradox''}: while complex, iteratively refined attacks are effective, they often trigger defensive compliance, whereas simple strategies achieve a higher rate of genuine persuasion (\textbf{84.7\%}). Our code and dataset are available at GitHub, https://github.com/cza1006/llm-persuasion-defense.