Certified automation reduces human effort in evaluating AI agents
Certified Selective Automation of LLM Agent Evaluation
Computation and Language
Summary
Evaluating how well AI agents perform tasks usually requires people to check their work, because automatic checks can make mistakes without warnings. The authors developed a method that lets an automatic system decide most cases while guaranteeing its errors stay under a set limit. They made a new way to measure this that works well even when tasks are similar and related, which was a problem for older methods. Their system can automatically verify about a third to half of evaluations on tools and web data, and can also help improve itself in new areas without needing new human labels.
What this means in practice
- •For ai evaluation teams: Reduce human workload by automatically certifying a large fraction of AI agent evaluations with guaranteed error bounds.
- •For machine learning engineers: Improve model performance in new domains by using certified automatic evaluation to safely generate training labels without human input.
Authors
Chengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni
Abstract
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectories stays below a budget alpha? Agent corpora resist the standard answer: many agents attempt the same tasks, so trajectories arrive in correlated clusters, and the i.i.d. certificates of existing selective-judging methods can overstate what is safe: a naive certificate can claim 98% automation while its realized error exceeds the budget in 17.5% of task resamples. We introduce a task-level bootstrap certificate that is valid in every regime we test while matching the naive certificate's coverage; finite-sample cluster-valid alternatives certify nothing at realistic task counts. Under this certificate, a 4B logprob judge trained with SFT and reject-weighted GRPO certifies 0.30-0.59 of evaluation on tool-use and web corpora at alpha=0.1, the only judge, among strongly elicited frontier models, certifying on both headline corpora. Certified coverage is predictable before training from base rate and discrimination alone (leave-one-corpus-out R^2=0.96). Finally, the certificate doubles as a self-training filter: pseudo-labels harvested inside certified regions have contamination bounded by alpha by construction (realized 0.000-0.041 across six harvests), letting a judge enter an unseen domain at in-domain strength with zero target training labels.