Provenance based runtime guard stops cascading attacks on llm agents
AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents
Cryptography and SecuritySoftware Engineering
Summary
Large language model (LLM) agents can cause harm by chaining normal actions that look safe on their own. The authors created AGATE, a tool that checks who allowed each action and where the data comes from to decide if the action is okay. It watches agent boundaries and records detailed evidence so decisions can be replayed for review. AGATE works with three existing LLM frameworks without changing their code and was tested on many attack and normal scenarios to verify its accuracy and limits.
What this means in practice
- •For security engineers: Enforce fine grained permission and data origin checks during large language model agent operations to prevent complex unauthorized actions.
- •For software platform developers: Integrate a unified authorization and provenance gate into existing agent harnesses to improve security without modifying host code.
- •For enterprise risk managers: Use provenance records and replayable decision evidence for forensic analysis of automated agent actions across business workflows.$Commercial implications: Enables audit and compliance tools for organizations deploying LLM automation with clear traceability and accountability.
Authors
Xiaorui Zhang, Zhuoran Cheng, Kailin Liu, Zhaoxi Sun, Shiyu Fan, Tongyu Yuan, Bin Yuan, Weizhong Qiang, Deqing Zou
Abstract
LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.