Policy ambiguity causes unreliable agent evaluation scores
Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error
Computation and Language
Summary
Agent benchmarks often assume that a set of rules (policies) clearly say what an agent should do in every situation. This paper shows that real policies written in natural language can be vague or contradictory, allowing multiple right answers that benchmarks don’t capture well. As a result, scores given to agents can fluctuate and be misleading. The authors analyze tasks where policy ambiguity mixes with tool limits, causing inconsistent agent behavior and evaluation. They suggest that carefully checking policies before judging agents can improve evaluation reliability.
What this means in practice
- •For benchmark developers: Detect ambiguous or contradictory policy statements before creating evaluation datasets to improve agent scoring accuracy.
- •For ai system evaluators: Identify tasks where policy complexity and tool limits cause unreliable scores to better interpret agent performance results.
Authors
Hongliu Cao
Abstract
Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory cannot capture. Auditing two $τ^2$-bench domains, we develop a taxonomy of such policy loopholes and show that affected tasks produce unreliable scores: they lower scores across different models in different ways and make every model less consistent across repeated trials. A cross-domain comparison reveals that exploitability requires both policy ambiguity and tool permissiveness: when policy complexity exceeds what tools can enforce, agents resolve gaps inconsistently and scores become unreliable. Policy specification quality sets the ceiling on evaluation quality. Benchmark developers should audit policies before collecting gold annotations.