FlowSeal stops privacy leaks in AI assistants using outside control
Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control
Cryptography and SecurityArtificial Intelligence
Summary
Personal AI assistants that use large language models (LLMs) often need to access sensitive user information, creating risks that private data might be leaked. The paper's authors show that existing privacy protections can be bypassed by clever tricks that confuse the AI inside the same conversation. They introduce FlowSeal, a new defense that controls data flow outside the AI’s internal context, using careful tracking of information and limits on what can be shared. Tests show FlowSeal greatly reduces leaks while still allowing useful tasks to be done, no matter which LLM powers the assistant.
What this means in practice
- •For ai assistant developers: Prevent unintended private data disclosure in AI assistants by enforcing data flow rules outside the language model context.
- •For enterprise security teams: Deploy external information flow controls to secure AI tools that handle sensitive corporate communications and data.
Authors
Minsun Shim, Ramisha Raida Karim, Ruthwik Jakkula, Kaiwen Zhou, Xin Liu, Xin Eric Wang, Zhou Li
Abstract
Personal AI agents built on large language models (LLMs) are increasingly given access to a user's private data and communications in order to provide personalized assistance. This access creates a persistent privacy risk: the agent must decide whether a given sensitive information should be disclosed to a particular party. Existing defenses address this by making the agent's backend LLM more privacy-preserving through stronger system prompts, training, or explicit consent-checking procedures, but this approach has a structural challenge: whenever enforcement is a judgment the LLM makes over the same conversational context an adversary controls, the enforcement mechanism and the attack surface coincide. We demonstrate this against existing defenses with three new attacks that require only ordinary agent interaction and no prompt injection: Collaborative Workspace Lure reframes an extraction attempt as collaborative work; Semantic Obfuscation Attack induces disclosure through omission rather than through anything the agent writes; and Channel Decoupling Attack splits the extraction request and the disclosure across independent channels. All three achieve substantially higher leak rates than the attacks these defenses were originally designed to withstand. Guided by this observation, we present FLOWSEAL, a defense that enforces confidentiality through a tool-level interceptor outside the LLM's context, grounded in data provenance and an information-flow-control lattice with controlled declassification. Evaluated across three benchmarks, five prompt-based baselines, and eight attacks, including a real agent executing live tool calls through MCP, FLOWSEAL reduces leak rates to near zero (e.g., 52.2% to 0.5% against Collaborative Workspace Lure) while preserving task utility, regardless of the underlying LLM backend.