LLM coding agents measured for security issues along code progress
Trajectory-Level Security Debt in LLM Coding Agents
Cryptography and SecuritySoftware Engineering
Summary
Software tools powered by large language models (LLMs) often create many versions of code before settling on a final answer. The authors found that just checking the finished code misses security risks that appeared and changed along the way. They propose a new way to track and add up potential security problems throughout these coding steps instead of only at the end. This measure, called the Security Debt Line Integral, helps understand how security issues develop as code improves but still needs testing for guiding fixes or confirming real dangers.
What this means in practice
- •For software security teams: Monitor security risks continuously during automated code generation to better understand evolving vulnerabilities.
- •For software development teams: Use trajectory-level security insights to evaluate and improve coding agents’ generated code beyond final test results.
Authors
Prateek Kumar Rajput, Abdoul Kader Kabore, Yewei Song, Melissa Tessa, Tailia Malloy, Jacques Klein, Tegawendé F. Bissyandé
Abstract
LLM coding agents can traverse hundreds of intermediate code states before submitting a solution. Evaluating only the final artifact leaves the evolution of security findings unmeasured. We introduce the Security Debt Line Integral (SDLI), which accumulates static-analysis risk when an agent reaches a new best test pass ratio. We instantiate it with four static application security testing (SAST) tools and study artifacts from 830 passing SWE-bench runs, 712 ProgramBench final workspaces, and 13 public MirrorCode trajectories. The two large populations use the final-state special case of SDLI. Two-tool Common Weakness Enumeration (CWE) class agreement occurs in 3.9% of SWE-bench runs and 26.2% of the 80 ProgramBench runs passing at least 90% of official tests. These are scanner findings, not validated vulnerability rates. Excluding three advisory-heavy classes reduces the latter rate to 6.2%. Same-task runs differ in their measured scores, while one reconstructed ProgramBench run exposes persistent findings from its first implementation write. A repair case study reduces the scanner signal while preserving tested behavior, but also reveals sensitivity to equivalent API rewrites. SDLI offers a way to study progress and security findings together. Its value for steering agents and confirming exploitable vulnerabilities remains to be established.