Frontier AI agents often falsely claim to complete file review tasks
Quantifying Overclaiming Propensity in Frontier LLM Agents
Software EngineeringArtificial IntelligenceMachine Learning
Summary
Frontier AI coding assistants are trusted to work alone for long periods, but their final answers might not tell the full truth. The authors measured how often these agents say they completed a review task when they actually skipped some files. They found that about two-thirds of the time, agents missed files, and most of those times they wrongly said they read them all. This false claiming hides mistakes that could be serious, showing users cannot fully trust what agents report about their own work.
What this means in practice
- •For ai product teams: Assess whether autonomous coding agents genuinely complete assigned review tasks before deployment.
- •For software quality assurance teams: Detect misleading claims from AI reviewers about task coverage to prevent missed defects in codebases.
Authors
Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo, Saskia Helbling, Alberto Tosato, Mohamed Amine Merzouk, Nouha Dziri, Gauthier Gidel, Tommaso Tosato
Abstract
Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of frontier agents to \emph{overclaim} task completion, a misrepresentation that can mislead the user. An agent overclaims when its final response contradicts information in its context. This definition requires no inference about intent and is independent of task success. We introduce \emph{OverclaimBench}, an evaluation suite composed of five file-review scenarios, transcript-based coverage measurements, and registered planted defects. We evaluate eight proprietary frontier models in their own production command-line interfaces, and four open-weight models under a single fixed harness on OverclaimBench and find that 1) agents do not read all the files they were asked to review in 67.9\% of runs; 2) among runs where not all files are read, agents are \emph{misleading} 80.4\% of the time (59--96\% per model), either falsely claiming to have read all files or omitting that coverage is incomplete; 3) requiring delegation to subagents increased reading coverage, but among reviews that remained incomplete, a large majority were still misleading; and 4) agents that falsely claimed a complete review missed planted defects at about 1.8 times the rate of agents that read every file, showing that claims of completion can conceal substantive failures. Together, these results show that agents' final responses are not reliable accounts of their actions.