Kernel data improves security detection of large language model agents

On the Effectiveness of Kernel-Level Evidence for Agent Security

Cryptography and SecurityArtificial Intelligence

Summary

Large language model (LLM) agents often have deep access to a computer’s system, but current security checks only look at their messages and commands, missing hidden threats. The authors studied both what the agents say and what their system calls do to catch suspicious behavior better. They created a large set of examples combining these two views and found that using both together detects attacks more accurately than using either alone. Their approach also works well on different types of attacks and in different agent setups. This shows that looking deeper inside the system helps improve security for these AI tools.

What this means in practice

  • For security teams: Detect malicious behavior in AI agents by combining application messages and low-level system call data for improved threat detection.
  • For infrastructure operators: Use cross-layer telemetry to monitor AI agents running on hosts with broad access, enhancing security against hidden kernel-level attacks.

Authors

Spencer King, Zhilu Zhang, Mikhail Kuznetsov, Kay Liu, Baris Coskun, Wei Ding

Abstract

LLM agents are deployed into infrastructure that grants them broad host authority, yet existing agent-security benchmarks and defenses operate almost exclusively at the application telemetry layer: the served tool manifest, the user prompt, and the model's messages. Some threats, however, smuggle malicious instructions and actions past the application boundary, leaving them invisible to that layer. In this work, we bridge that gap by pairing application-level agent telemetry with kernel-level syscall traces to present the first paired-evidence characterization of kernel-level versus application-layer signal for agent security. To quantify the value of the enhanced telemetry, we introduce Agent Cross-Layer Evidence (ACE), a paired-session corpus of 4,047 sessions and 17 threat models spanning six delivery-vector families and 14 of the 25 OWASP LLM and agentic threat categories, organized into 12 attack mechanics with per-mechanic characterization of where the most discriminative evidence lies. Across four distinct detector families, we find that kernel evidence is discriminative on its own and that composing it with application-layer evidence generally outperforms either single-layer view, revealing complementary signals that single-layer analyses can miss. We further demonstrate generalization to unseen attack families and transfer to an alternate agent runtime. Together, these findings establish the value of cross-layer evidence for agent security.