Selection-Aware Stress Testing for Interactive Agents

2026-08-31Machine Learning

Machine Learning
AI summary

The authors address a problem where researchers pick the best way to do a task based on one set of data and then use the same data to find weaknesses, which can bias results. They propose a new method called Selection-Aware Semantic Stress Testing (SASST), which separates data into discovery and confirmation sets to fairly check if claimed advantages hold up. Their method includes statistical checks to avoid false claims and can sometimes conclude that no solid claim can be made. In tests, SASST showed that initial improvements often disappear when confirmed with new data, highlighting the importance of this cautious approach.

agent evaluationbenchmarkingworkflow selectionstress testingstatistical confirmationtask reweightingcluster assumptionsBonferroni correctionGaussian coveragepaired comparison
Authors
Yang Xu, Chenang Li, Jiefu Zhang, Haixiang Sun, Zhou Li, Vaneet Aggarwal
Abstract
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $τ$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.