AgentHop benchmark reveals large language model multi-step answer challenges

AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

Computation and LanguageArtificial Intelligence

Summary

Some computer programs called large language models try to answer complicated science questions by taking multiple steps and using tools. The authors of this paper created a new test called AgentHop to find out exactly where these programs make mistakes, such as searching for information, combining facts, using tools, or managing resources. They studied 19 different models and found that these models make different kinds of mistakes and have unique behaviors. AgentHop helps better understand how these models work so they can be improved in the future.

What this means in practice

  • For ai developers: Use AgentHop to diagnose and improve failure modes in multi-step question answering AI systems.
  • For ai benchmark designers: Incorporate AgentHop’s diagnostic axes into new benchmarks to get detailed insights on AI agent behavior.

Authors

Chanhee Park, Jeongho Yoon, Sungbin Han, Hyeonseok Moon, Heuiseok Lim

Abstract

Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.