Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs

Artificial Intelligence

Summary

The authors studied how multi-agent systems using large language models (LLMs) share information by summarizing previous interactions into a handoff message. They found that these summaries tend to keep important facts but lose the details about rules or boundaries for how that information should be used, a problem they call "summary collapse." This loss leads to unintended privacy leaks. Their experiments showed that being explicit about boundaries greatly reduces leakage, and using an audience allowlist works best to prevent leaks. Simple fixes like redacting text or changing prompts only help a little.

Authors

Yian Wang, Agam Goyal, Eshwar Chandrasekharan, Hari Sundaram

Abstract

Multi-agent LLM systems often coordinate by compressing an upstream interaction into a handoff artifact that downstream agents treat as shared state. We show that this handoff step is a structural source of privacy leakage: summaries preferentially preserve operational facts while weakening the boundary metadata that governs how those facts may be used---a failure mode we call \emph{summary collapse}. On a controlled multi-agent coordination testbed we measure marker survival with a human-validated judge ($κ= 0.74$), where $σ_b = 1$ means every boundary marker survives verbatim and $σ_b = 0$ means all are lost. Boundary-marker and operational-fact survival are nearly uncorrelated at the handoff level on both GPT-5-mini and DeepSeek-R1-32B (Pearson $r$ near zero): uncompressed free-text handoffs preserve boundaries at $σ_b \approx 0.80$, whereas a $25$-word budget drops $σ_b$ to ${\approx}0.57$ while operational-fact survival stays near ceiling. Controlled downstream tests reveal that protection depends on \emph{boundary explicitness}: vague languages leak in $73\%$ of GPT and $50\%$ of DeepSeek cases, while explicit constraints reduce leakage to under $15\%$ across all three tested models. A no-handoff single-agent control further shows the failure is not reducible to multi-agent topology as direct full-marker access still leaks more often than the operationalized handoff. Prompt-only mitigation and exact-string redaction only partially address the problem, while a gold-derived audience allowlist nearly eliminates leakage across models, showing that correctly identifying audience boundaries is the key factor.