Agents must keep and use key information to complete complex tasks

Dude, Where's My State? Execution Information Requirements for Stateful Agents

Artificial Intelligence

Summary

Computers that perform long tasks need to remember important information from earlier steps to succeed later. The authors propose a way to figure out the minimum information that must be kept accessible, called Execution Information Requirement (EIR). They also created tools to test how agents handle keeping, losing, or recovering this information. Their findings show that just having enough storage isn't enough—agents must also keep the right information, avoid errors, and finish recovery steps. This helps people understand if agents truly manage the information they need to get tasks done correctly.

What this means in practice

  • For software developers: Design and evaluate software agents ensuring they preserve and recover required task data to avoid failures during long-running processes.
  • For automated system testers: Use the framework to identify when system agents lose essential information, improving detection of failure causes in complex workflows.

Authors

Nikita Mehrotra, Ashish Tiwari, Priyanshu Gupta, Sumit Gulwani

Abstract

Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.