Coding agents often duplicate code despite functional success
Do Coding Agents Reuse Existing Code or Reinvent the Wheel?
Software EngineeringArtificial IntelligenceComputation and Language
Summary
Coding agents that write software sometimes copy the same code multiple times instead of reusing existing code. The authors built a benchmark called RepoReuse to check if these agents reuse code or reinvent it during multiple steps of software development. They found that agents frequently leave duplicated code even when their code passes all tests, which could cause extra maintenance work. This shows that just checking if code works well is not enough to evaluate coding agents.
What this means in practice
- •For software engineering teams: Assess code duplication in multi-step coding tasks to improve maintenance and audit processes.
- •For devops teams: Monitor automated code generation pipelines to detect redundant code and reduce technical debt.
Authors
Dongsheng Ma, Sizhe Wang, Xinyi Huang, Zhengren Wang, Yuhan Wang, Luyang Si, Xincheng Wei, Wentao Zhang
Abstract
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8\% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.