Fixed language model judges make version-dependent mistakes evaluating agents

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

Machine Learning

Summary

Evaluating improvements in AI agents using language-model judges that are kept the same over time can lead to errors that depend on which versions of the agents are compared. The authors studied multiple coding and customer-service AI agents and found that the judges often disagree depending on agent versions, sometimes wrongly approving inferior upgrades. They suggest that relying solely on these fixed judges or old calibration can cause significant mistakes. Instead, they recommend comparing agent outputs directly with explicit reference standards and current data audits.

What this means in practice

  • For software engineering teams: Improve evaluation reliability of coding agents by using paired audits and explicit references rather than fixed language-model judges alone.
  • For customer support automation teams: Better assess upgrades to customer-service AI by comparing agent outputs with current references and avoiding outdated judge calibrations.

Authors

Jiapeng Li

Abstract

Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned agent-task cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare improvements that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge confidently reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and reference outcome, while a task-solvability prediction from AgentRewardBench reverses sign in SWE-bench. Transporting old-version calibration raises mean absolute comparison error on SWE-bench from 3.8 to 19.5 percentage points; 24.6% of ratio-bootstrap draws are undefined near the correction boundary. A tuned paired audit narrows a classical interval by only about 5% at 80 labeled tasks. A randomized three-arm test does not support the predicted increase in false acceptance from showing the agent's final report (all three Holm-adjusted p-values = 1.0). These results favor explicit reference standards and paired audits of current outputs over judge-only release decisions or transported old-version calibration.