Papers for

automated quality assurance teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Paired evaluations improve testing for visual recognition errors

What Paired Evaluations Reveal under Visual Perturbations

Abstract: Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We investigate how this image correspondence supports two needs in robustness evaluation: interpreting paired evaluation results and prioritizing samples for physical testing. To interpret paired evaluation results, we fix both sets of prediction records and vary their correspondence within each class. We prove that classwise correct-correct counts give the same sharp bounds on lost acceptance and mean true-class probability decrease among retained-correct inputs as any feasible five-state refinement. Distinguishing persistent from changed wrong answers can further constrain accepted-error transitions, while shared correspondence can establish policy orderings left unresolved by separate cost intervals. To prioritize samples for physical testing, we retain each image's synthetic responses and rank clean-correct images by their mean true-class probability under corruption. Across 44 classifiers, testing the highest-risk 20% finds 67% and 45% of failures under mild screen and print recaptures, versus 58% and 36% for clean confidence and 60% and 37% for an equal-size natural-transformation average. With both probability averaging and an A3Rank scoring adaptation, the tested corruption set yields higher mean failure recall than the natural-transform set; differences between scores depend on the source and budget. Together, these findings show that the value of correspondence depends on the evaluation objective: classwise counts suffice for specified reliability bounds, while image-specific synthetic responses improve the allocation of physical tests within the evaluated pool.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Robustness tests check how well image recognition systems work when pictures are changed or disturbed, but real-world checks are expensive. The authors show that comparing predictions on original and changed images together reveals more about errors. They prove that simple counts of matching correct answers can reliably estimate how systems lose confidence, while using detailed comparisons helps focus on which images to test physically. This approach finds more mistakes using fewer physical tests than traditional confidence scores or averaging methods.
Open → 2609.35583v1

Agent harnesses improve task success and reduce errors in AI planning

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Abstract: Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

Thu 17 SeptArtificial Intelligence
The gist
Planning tools help AI agents organize tasks, check progress, and decide what to do next. This paper shows that giving AI agents carefully planned instructions boosts their success, especially for harder tasks. Checking answers after tasks finish helps avoid wrong approvals but can sometimes hold back good results. Depending on how costly mistakes are, either detailed planning or strict checking is more valuable. The authors also show that just checking results can catch most errors cheaply without the full planning setup.
Open → 2609.20474v1