Ai agents can spread harmful content through shared memory artifacts

Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents

Artificial IntelligenceComputation and LanguageCryptography and SecurityMachine Learning

Summary

Large language model assistants can keep memories and share files that help them work with each other. The authors found that bad content can sneak into these shared files and then be passed along from one assistant to another, like a virus. This attack can last a long time and reach many different assistants, even if they are supposed to be separate. The researchers tested this across different setups and saw that some attacks spread widely and lasted for many steps.

What this means in practice

  • For ai security teams: Detect and prevent attacks that spread through shared files and memories between AI assistants.
  • For platform operators: Monitor and control information flows in AI assistant ecosystems to limit the spread of harmful content via persistent artifacts.

Authors

Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina, Ivaxi Sheth, Mario Fritz

Abstract

Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.