Better documentation does not improve coding agent fixes in repositories
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Software EngineeringArtificial IntelligenceComputation and Language
Summary
People often think that good natural-language explanations of code help automated coding assistants fix problems better. The authors built a way to test how well code descriptions work by seeing if code rewritten from them still passes tests. They found shorter or longer descriptions don’t matter as much as how complete the explanations are. However, when they tried using improved descriptions to help coding agents fix real issues in software projects, the descriptions did not help. This suggests that having the code and the problem alone is usually enough for these agents.
What this means in practice
- •For software engineering teams: Use benchmark tools to measure if code descriptions truly capture code behavior during documentation creation.
- •For code quality assurance teams: Evaluate documentation completeness to ensure software descriptions accurately reflect code functionality for testing purposes.
Authors
Md Shohel Arman, Igor Molybog
Abstract
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.