Agent comparisons often overcount independent evidence in evaluations

When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate

Artificial Intelligence

Summary

When comparing many agents by looking at every pair, it might seem like there is a lot of independent information. But because some agents appear in multiple pairs, the actual unique evidence is less than the number of pairs suggests. The authors show that if agents are seen as fixed, pairs summarize them exactly, but if agents are random samples from a larger group, uncertainty depends more on the number of agents than pairs. Their findings help clarify how to interpret comparison results, especially in leaderboards that rate agent performance.

What this means in practice

Authors

Wei-Jung Huang

Abstract

When an agent benchmark compares every pair of leaderboard entries, the number of comparisons can look much larger than the independent evidence behind them: A versus B and A versus C both reuse A. Whether this reuse affects inference depends on what the analysis is meant to describe. If the board and its outcomes are fixed, the all-pairs mean is an exact summary of those entries, and any interval must come from another declared source of randomness. If the entries are instead treated as iid draws from a population of future configurations and the pair rule is regular and nondegenerate, the same mean is an order-two U-statistic whose first-order uncertainty depends on the number of configurations, not the number of pairs. We use near ties as the running example, but the distinction extends to other symmetric pair summaries when their regularity conditions hold. We examine both interpretations using a fixed SWE-bench Verified snapshot and an exact binary model with known truth. On SWE-bench, intervals that accounted for shared configurations were more than twice as wide as a pair-iid reference that treated the pairs as independent. In the exact model, pair-iid coverage fell far below the nominal level when edges shared endpoints but remained near nominal for matched independent edges. Results on two other fixed leaderboards show that exact summaries also depend on which pairs are included and how they are weighted. An all-pairs analysis must therefore state what is fixed, what is sampled, and how it handles shared entries and pair aggregation.