Harness design shapes software engineering agent performance
Beyond the Model: Demystifying Harness Effects in Software Engineering Agents
Software Engineering
Summary
Software agents powered by large language models do coding tasks, but how well they work depends not just on the base model, but also on the 'harness'—the system that manages their interactions and actions. The authors studied different harness designs and found that more complex harnesses help mainly with harder coding tasks and stronger models. They also pinpointed which harness parts boost performance and which can hurt it. Their work shows building better harnesses is key to improving coding agents, beyond just improving the language model itself.
What this means in practice
- •For software engineering teams: Improve automated code repair and repository-level tasks by selecting or designing agent harnesses suited to model strength and task complexity.
- •For cloud platform developers: Build more effective AI coding assistants by integrating modular harness components like structured tool use and subagents into agent pipelines.
Authors
Haichuan Hu, Quanjun Zhang, Shengcheng Yu, Zhifei Chen, Tianyu Luo, Chunrong Fang, Zhenyu Chen, Liang Xiao
Abstract
Large Language Model (LLM)-based agents are increasingly used for software engineering tasks, yet their performance is not determined by the base model alone. The agent harness substantially shapes how SE agents interact with repositories, execute actions, and validate solutions. However, the role of harness design remains insufficiently understood, especially across different models, tasks, and harness components. In this paper, we present a systematic empirical study of harness effects in SE agents. We first evaluate two representative harnesses, mini-SWE-agent and OpenCode, with ten models from two prominent open-weight model families, Qwen and DeepSeek, on three benchmarks: SWE-bench Pro, ProgramBench, and GitTaskBench. We then construct NanoHarness, a lightweight modular harness built on top of mini-SWE-agent, and use it to analyze five representative harness components: tool registry, context compression, explicit planning, subagents, and lazy skills. Experimental results show that harness effectiveness depends jointly on model capability and task type. Complex harnesses provide diminishing marginal gains on SWE-style issue repair as model capability improves, but can benefit stronger models on more complex and open-ended repository-level tasks. Component-level analysis on ProgramBench further shows that structured tool use and task-specific subagents provide the most stable improvements, while context compression and general subagents can hurt repository-generation performance. When combined, NanoHarness improves over mini-SWE-agent by 7.37 and 6.21 percentage points on Qwen3.7-Max and DeepSeek-V4-Pro, respectively, recovering most of the gains of product-level harnesses. These findings highlight harness design as a first-class factor in SE-agent performance and provide insights for building more effective and efficient coding agents.