Large language models can accidentally reveal removed sensitive information

"Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables

Cryptography and SecurityArtificial Intelligence

Summary

Sometimes, when people use AI chatbots to help write things, they ask the model to remove sensitive info like passwords before sharing. But the AI might mention that it removed something in a way that reveals the secret anyway. This paper studies when and how often this happens and shows it’s a real risk with many popular AI tools. The authors also test ways to stop the AI from accidentally sharing secrets while still keeping the important parts of the message.

What this means in practice

  • For software developers: Detect and filter AI-generated comments that reveal removed sensitive data before sharing documents externally.
  • For security teams: Implement safeguards against accidental leaks from AI assistants when generating or revising confidential content.

Authors

Yage Zhang, Yukun Jiang, Yang Zhang

Abstract

Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such statements revision traces. For example, after a user removes the password before sharing a configuration file, the model may delete it but leave a comment saying, "Removed the password 'No****4!' as requested." A third-party recipient who sees only the delivered file can therefore recover the withdrawn password from the comment. In an in-the-wild analysis of three public conversation corpora, we identify 26,753 revision requests, of which 2,363 (8.8%) leave revision traces. We study them in greater depth under controlled conditions by introducing RevLeakBench, a benchmark of 100 tasks across five scenarios with a conversation track and an agent track. We measure trace occurrence, withdrawn-item recovery, trace position, and required-content retention. Across six models, about half of the deliverables in both tracks state the edit after a revocation, and a reader that sees only the deliverable can recover the withdrawn item from about 13% of them. Telling the model that its entire reply will be forwarded to the recipient still leaves revision traces in 36.4% of the deliverables. We compare prompt defenses and a delivery boundary, and propose an output-side filter that sharply reduces recovery with little loss of required content. We believe our work can benefit efforts to understand and mitigate unintended disclosure in LLM interactions.