Papers for

automated system testers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Safety and recovery improve planning for household AI agents

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

Abstract: Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be separated from another. We present SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate (~250 lines of Python, zero tokens, $O(|π|)$) that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact; a hybrid seed+live memory store supports cold-start coverage. We evaluate under a leak-free protocol (leave-one-out retrieval) over five open-weight models and a 75-task AI2-THOR benchmark. On the standard benchmark goal-completeness saturates (52% of instances trivially solved) and SAGE ties strong hierarchical baselines. On a harder, method-agnostic multi-goal composition, SAGE's completeness lead re-emerges large (+0.06 to +0.23 across four models). Under injected mid-execution failures, SAGE recovers as reliably as whole-plan replanners at 2.4-3.3x fewer LLM calls. As a verify-before-execute gate, the symbolic monitor blocks unsafe actions before actuation and raises simulator-reported step-success for every planner tested (up to +0.11), a signal the verifier never sees (non-circular). Because the gate calls no model (0.008 ms/plan), it is a safety layer that runs essentially free on the edge: SAGE planning reproduces its quality on a Jetson AGX Orin, where small-model verification helps most. We release the benchmark, the leak-free protocol, the recovery and safety-gate harnesses, and a verifier-portability study (auto-induced on ALFWorld, 0.89 held-out).

Mon 28 SeptRoboticsArtificial Intelligence
The gist
Household robots that use language models often plan actions without checking if those actions really make sense in their environment, which can cause mistakes. The authors present SAGE, a method that adds a simple safety check to block impossible actions and edits plans locally when something goes wrong. This makes planning safer and more efficient, especially in complex tasks or when unexpected errors happen during execution. Their tests show SAGE helps robots recover from failures more reliably while using fewer calls to the language model.
Open → 2609.34268v1

Agents must keep and use key information to complete complex tasks

Dude, Where's My State? Execution Information Requirements for Stateful Agents

Abstract: Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.

Sat 26 SeptArtificial Intelligence
The gist
Computers that perform long tasks need to remember important information from earlier steps to succeed later. The authors propose a way to figure out the minimum information that must be kept accessible, called Execution Information Requirement (EIR). They also created tools to test how agents handle keeping, losing, or recovering this information. Their findings show that just having enough storage isn't enough—agents must also keep the right information, avoid errors, and finish recovery steps. This helps people understand if agents truly manage the information they need to get tasks done correctly.
Open → 2609.32687v1