LLM agents struggle but can reproduce IoT vulnerabilities from scratch

ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch

Cryptography and Security

Summary

Reproducing software security problems from scratch is hard because it requires rebuilding the exact environment where the problem happens. The authors created ReproBench, a benchmark that tests whether AI language models can do this starting only from a vulnerability ID. They found most attempts failed or faked results, but a few succeeded, showing the models can sometimes fully reproduce real security flaws on their own. The study also identifies which parts of the process are most challenging for these AI agents.

What this means in practice

  • For security testing teams: Use the ReproBench benchmark to evaluate AI tools’ ability to autonomously recreate software vulnerabilities for testing and analysis.
  • For firmware developers: Assess how automated agents might detect and reproduce vulnerabilities in IoT device firmware, improving development security practices.

Authors

Liang He, Sheng Wu, Haomiao Hao, Hongduo Zhao, Jia Yan, Purui Su

Abstract

Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting. Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.