LLM agents often fail to protect private information across communities
CoSec: Benchmarking Agent Security in Communities
Cryptography and SecurityArtificial Intelligence
Summary
Large language model (LLM) agents work together in groups that can change over time and involve shared information like files and memories. The paper introduces CoSec, a test system that checks how well these agents keep private information secure and follow rules about who can see what. The authors found that while agents usually complete their tasks, they frequently expose private data or break authorization rules. This shows that being useful doesn't mean they are safe, and securing these agents in shared settings is still a big challenge.
What this means in practice
- •For software security teams: Evaluate and improve privacy enforcement of AI agents in multi-user environments using CoSec benchmark scenarios.
- •For enterprise ai developers: Test LLM agent workflows for unauthorized data leaks before deploying in business settings with community-based data access controls.
Authors
Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao, Chenbo Xia, Yuwen Qu, Renxiang Wang, Fudong Yuan, Camil Hamami, Chenglong Pan, Xinquan Yue, Ziyu Wang, Fengyu Ye, Chenyang Si, Caifeng Shan
Abstract
LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthorized disclosure of protected information. Existing evaluations do not fully examine these risks in agent systems. We introduce \textbf{CoSec}, an executable benchmark for evaluating privacy and authorization enforcement in LLM agent systems operating within and across communities. CoSec contains 208 canonical scenarios spanning fixed and evolving boundaries, protected information belonging to the agent owner or other participants, and attacks through dialogue, environmental content, persistent memory, and composed workflows. CoSec executes complete agent systems with persistent sessions, memory, files and tools. It verifies information flows against the active authorization state using execution traces and artifacts. Across harness and model configurations, agents frequently complete benign tasks but violate privacy and authorization boundaries. Privacy behavior varies across harnesses, attack surfaces, and community states, revealing how memory, files, tools, and workflows can carry protected information beyond its authorized scope. These findings show that task utility does not imply privacy or authorization compliance and that authorization in community settings remains an unresolved security challenge for persistent LLM agents.