Corrected agent experiences improve task performance but with limits
After the Fix: How Corrected Agent Histories Transfer to Related Tasks
Artificial IntelligenceComputation and LanguageSoftware Engineering
Summary
When a computer agent fixes a mistake it made in one task, it can sometimes do better on a related task next time. The researchers tested thousands of cases where agents tried tasks, failed, then corrected their approach before trying again. They found that repaired experiences helped more in some setups than others, but improvements depended on comparing corrections to both old attempts and new fresh starts. Sometimes better fixed memory didn't come from better corrections but from worse initial attempts, showing the challenge in fully benefiting from corrections over separate tries.
What this means in practice
- •For autonomous system developers: Improve decision-making in AI agents by incorporating corrected past experiences when tackling related tasks to boost success rates.
- •For robotics engineers: Optimize robot task workflows by selectively reusing corrected execution traces to handle variations of previously failed tasks more effectively.
Authors
Yanfei Zhang, Xu Lin
Abstract
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-state inheritance, under eleven conditions. ThinkingBox's Full/Skill/Hybrid correction gains are 44/29/32 percentage points, with corrected performance 25/22/18 points above independence; inference weakens at the task-family level. Yet 12 of Full's 15-point larger correction gap over Skill come from worse uncorrected performance, not better corrected memory. Moreover, 22 of Full's 46 upward transitions restore observed baseline success. Neither APEX regime establishes comparable aggregate correction benefits. Action evidence connects workflow gains with reusable obligations and convention conflicts with source-local choices. Text APEX's accepted execution reaches 52% versus its summary's 40%, without robust global/group-level superiority or an estab- lished advantage over independence. Smaller handoffs reduce input but increase calls. The value of repairing experience is therefore distinct from the value of reusing it: memory updates require both a previous-version reference and a fresh-start reference.