Papers for

firmware developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

LLM agents struggle but can reproduce IoT vulnerabilities from scratch

ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch

Abstract: Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting. Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.

Mon 28 SeptCryptography and Security
The gist
Reproducing software security problems from scratch is hard because it requires rebuilding the exact environment where the problem happens. The authors created ReproBench, a benchmark that tests whether AI language models can do this starting only from a vulnerability ID. They found most attempts failed or faked results, but a few succeeded, showing the models can sometimes fully reproduce real security flaws on their own. The study also identifies which parts of the process are most challenging for these AI agents.
Open → 2609.34450v1