Continual search improves AI failure diagnosis in long tasks

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Artificial IntelligenceHuman-Computer InteractionMachine LearningSoftware Engineering

Summary

When AI systems work on long and complex tasks, they generate huge records of their actions, making it hard to figure out why they fail. The authors found that current methods using large language models (LLMs) have trouble pinpointing the true causes because the clues are hidden across many steps. They created a new approach called Continual Search that repeatedly prompts the AI to keep looking for important evidence instead of stopping too soon. This method significantly boosts the accuracy of diagnosing failures, even helping smaller models outperform larger ones when guided properly.

What this means in practice

  • For ai system engineers: Diagnose failures in long-running AI tasks by iteratively searching execution logs to find root causes more accurately.
  • For software reliability teams: Use continual search to improve automated failure analysis and reduce manual review in complex system logs.

Authors

Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta

Abstract

The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execution traces grow larger. They struggle because relevant information is often sparse, distributed across distant actions, and disconnected from the visible failure, reducing root-cause attribution to a massive search problem. Existing RCA methods typically rely on one-shot LLM judgments to diagnose failures from execution traces. While effective for shorter trajectories, these judges tend to settle on a plausible diagnosis early, leaving critical evidence in longer traces unexamined. We introduce Continual Search, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence. We evaluate Continual Search across four existing RCA benchmarks. Recognizing the lack of massive execution traces in current benchmarks, we introduce MegaRCA-Mix to evaluate RCA at scale. MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks. Across multiple benchmark suites and model families, Continual Search consistently improves attribution performance. On MegaRCA-Mix, for example, it improves GPT-5.5's F1 score by more than 40\%, from $0.349$ to $0.498$. More interestingly, within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.