Leaderboard ranks for AI agents may not show clear winner

What Does an LLM-Agent Leaderboard Rank Actually Compare?

Artificial Intelligence

Summary

Leaderboards try to show which AI agent is better by ranking them, but these ranks don’t always tell the full story. The authors explain that differences in tasks, how results are judged, and scoring rules can make rankings unreliable for declaring one agent truly better than another. They propose a careful method to compare agents by checking the details and including uncertainty in the comparison. Their tests show that close ranks often don’t mean a clear winner, and different judging methods can change which agent looks best. Overall, a leaderboard score sums up some evaluation, but careful interpretation is needed to claim one agent is superior.

LLM agentsleaderboardevaluation metricspairwise comparisonuncertaintytask mixturelabel sourceutility rulesfinite-sample conditions

Authors

Wei-Jung Huang

Abstract

An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.