Papers for

ai benchmark designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AgentHop benchmark reveals large language model multi-step answer challenges

AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

Abstract: Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.

Mon 28 SeptComputation and LanguageArtificial Intelligence
The gist
Some computer programs called large language models try to answer complicated science questions by taking multiple steps and using tools. The authors of this paper created a new test called AgentHop to find out exactly where these programs make mistakes, such as searching for information, combining facts, using tools, or managing resources. They studied 19 different models and found that these models make different kinds of mistakes and have unique behaviors. AgentHop helps better understand how these models work so they can be improved in the future.
Open → 2609.34428v1

Benchmarks for AI often fail to measure intended skills reliably

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Abstract: Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, applying it to 56 capability and safety benchmarks across 53 models. We label benchmarks with substantively similar purported concepts to a shared assigned concept, and ask whether model rankings on benchmarks with the same assigned concept correlate more strongly than rankings on benchmarks with different assigned concepts. We ask analogous questions at the item level using item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting these concepts may be conceptualized inconsistently across benchmarks. For assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting these capability concepts may not discriminate well from one another. In some cases, benchmarks that share design elements (e.g., score format) correlate more strongly than benchmarks with the same assigned concept. Finally, some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with benchmarks sharing their own assigned concept, suggesting they may measure a different concept than they purport to. For example, BBQ-accuracy correlates more strongly with benchmarks labeled reasoning than with benchmarks that share its assigned concept, bias. To support future empirical work on benchmark validity, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.

Tue 8 SeptComputers and Society
The gist
AI benchmarks are tests used to measure specific abilities of AI models, like reasoning or fairness. The authors found many benchmarks don’t consistently measure the skills they claim to, meaning results can be confusing or misleading. Sometimes benchmarks measuring different abilities show similar results, making it hard to tell skills apart. This study highlights problems in how AI abilities are tested and shares data for improving future evaluations.
Open → 2609.08812v1