Agent harnesses improve task success and reduce errors in AI planning

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Artificial Intelligence

Summary

Planning tools help AI agents organize tasks, check progress, and decide what to do next. This paper shows that giving AI agents carefully planned instructions boosts their success, especially for harder tasks. Checking answers after tasks finish helps avoid wrong approvals but can sometimes hold back good results. Depending on how costly mistakes are, either detailed planning or strict checking is more valuable. The authors also show that just checking results can catch most errors cheaply without the full planning setup.

What this means in practice

Authors

Yukun Zhang, Kemu Xu, Yishen Chen

Abstract

Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.