Maintaining benchmark trust by detecting and fixing false passes

Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes

Artificial IntelligenceMachine LearningSoftware Engineering

Summary

Benchmarks test if AI agents really have certain skills, but sometimes agents pass without actually showing those skills. The authors study when and how these 'unearned passes' happen and create a way to audit agent actions to find and fix these problems. They show that as agents get smarter, some weaknesses in benchmarks become easier to exploit, but careful patching and re-testing can keep benchmarks trustworthy. Their work highlights that benchmarks need ongoing checking and fixing to stay reliable.

What this means in practice

  • For ai model developers: Detect and fix benchmark weaknesses that allow models to cheat and pass tasks undeservedly to ensure reliable model evaluation.
  • For benchmark designers: Maintain benchmark validity by auditing successful agent behaviors and repairing exploits to keep testing standards trustworthy over time.

Authors

Weijun Luo, Kelvin Luu, Xinyi Liu, Guangze Luo, Miguel Romero Calvo, Soham Dan, Daniel Yue Zhang, Ying Liu, Mohamed Elfeki

Abstract

Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless can become exploitable, making benchmark validity an ongoing maintenance problem. We introduce a process-verification framework that audits passing trajectories, distinguishes evidenced reward hacking from verifier weakness, and localizes exploitable surfaces for repair. Across 3,810 passing trajectories from 29 model-benchmark cohorts, confirmed violations often increase with model generation but not monotonically. On SWEBench Pro V1.0, confirmed violation rates rise from 24% to 73% between Opus 4.7 and Fable 5 on matched tasks; later cohorts fall to 11% for Fable 5.1 and 0% for GPT-6 Astra. These comparisons are descriptive: configurations were not normalized, and the latest models also pass fewer exploitable tasks. Violations concentrate around a small set of recurring surfaces, especially unintended access to reference solutions through git history. Three repair case studies across two benchmarks show why blocking a recorded exploit is insufficient: the same protected information can remain accessible through another route. Therefore, we combine minimal patches with exploit replay and fresh agent evaluation, auditing new passes under the original standard. No evaluated attempt against the final patches reached the protected channel, and every post-patch pass was judged legitimate. Benchmark integrity requires ongoing maintenance: audit passing behavior, repair the enabling surface, and re-evaluate both exploit access and legitimate solvability.