LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

Cryptography and SecurityArtificial Intelligence

Summary

The authors reviewed research on AI agents that help automate software and systems security work. They found these agents can perform tasks and plan steps but lack clear boundaries for decision-making authority and ways to track their actions. The study organizes current methods, applications, and evaluation techniques from 2023 to 2026 to better understand progress and challenges in the field. Their work highlights the need for safer, more controllable, and auditable AI security agents in the future.

Authors

Jingjing Nie, Jiawei Guo, Krishna Meda, Haipeng Cai

Abstract

Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypotheses, invoke tools, interpret outputs, and revise plans. Large language model (LLM)-based agents, which can plan, use tools, retain state, and revise actions across multi-step workflows, are being rapidly adopted to automate this work. Given the consequences of delegating security decisions to autonomous systems, understanding how such agents are built, used, and assessed is crucial. Yet to this date, there remains a lack of systematic understanding of what has been done and how far we are in this field: the term "agent" is applied inconsistently, applications differ sharply in risk, and assessment protocols are often incomparable. To gain a comprehensive and coherent view of this area hence inform relevant future research, this paper provides a systematic literature review of the (1) technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, (2) applications, with respect to the security tasks served, and (3) assessment, including the datasets, outcome and trajectory metrics, safety measures, and baselines considered, over the peer-reviewed literature spanning the emergence of this area (2023--2026). Our synthesis reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable. In addition to knowledge systematization, we also extend our insights into the limitations of and challenges faced by current approach, application, and assessment designs, which shed light on potentially promising future research directions.