AI security incident reports need new details and protections

Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents

Cryptography and SecurityArtificial Intelligence

Summary

AI agents are computer programs that can act on their own, but they can be attacked in ways that are different from usual software. The paper looks at how to report security problems with these AI agents by talking to experts. It suggests including new types of information in reports, like what the AI remembers or what tools it uses. The authors also point out challenges such as keeping report data safe from leaks or attacks.

What this means in practice

  • For it security teams: Adapt incident reporting protocols to include AI agent behaviors and memory for clearer attack understanding.
  • For cloud service operators: Strengthen protective measures for incident report storage to prevent exploitation by attackers targeting reporting systems.

A position paper. It proposes an approach and reports no results.

Authors

Anastasia Pustozerova, Eugene Bagdasarian, Luca Beurer-Kellner, Battista Biggio, Nico Ebert, David Filip, Marc Fischer, Heather Frase, David Hofer, Juliane Hoffmann, Daphne Ippolito, Somesh Jha, Sean McGregor, Esfandiar Mohammadi, Luca Nannini, Cristina Nita-Rotaru, Alina Oprea, Kevin Paeth, Andrew Paverd, Jonathan Petit, Andreas Rauber, Christian Riess, John Sotiropoulos, Andreas Wespi, Kathrin Grosse

Abstract

AI agents are being deployed rapidly, accompanied by a growing number of AI-specific attacks and corresponding incidents. As incident reporting becomes increasingly important for legal compliance, governance, accountability, and security; current frameworks must be adapted to the unique characteristics of AI agents. In this paper, two editorial authors compare AI systems and AI agents and, drawing on input from 23 experts in academia and industry, identify the information required for reporting incidents where the security of AI agents is harmed. %involving AI agents. Potential reporting elements include, for example, agent memory and memory accesses, actual and potential levels of autonomy, and tool usage. Based on these findings, we identify several open research questions, including how to efficiently record incidents and how to determine whether vulnerabilities and incidents generalize. Expert feedback also highlighted potential reporting weaknesses, such as risks of data leakage and attacks targeting the reporting infrastructure itself, creating additional research needs. Lastly, we summarize privacy requirements and outline research directions for the secure and trustworthy deployment of AI agents.