Papers for

automation platform engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Tool agents improve request fulfillment with faster repair searches

One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair

Abstract: Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both operation choices and their concrete realizations, making complete-sequence regeneration costly. Moreover, regeneration repeats operation selection even when failure arises from how those operations are realized. The resulting challenge is to reduce this repetition while preserving exploration of alternative operations and realizations. Therefore, we formulate repair as hierarchical search over operation supports, which we introduce as sets of permitted operation types that define reusable search regions for concrete tool-call sequences. We propose ReCommit, a training-free, diffusion-guided framework for improving tool-agent failure recovery while reducing repair computation. ReCommit amortizes operation-level proposal computation across repair trials by reusing operation-type scores from a single parallel readout of a masked diffusion language model. These scores guide search across supports, while realization search explores alternative entity bindings, arguments, and action composition within each support. Experiments on real failures across four enterprise services in the Agent-Diff benchmark show 75.9\% and 63.2\% relative recovery gains with 61.3\% and 51.3\% reductions in mean full-budget repair time at repair budgets $B=3$ and $B=13$, respectively, over the strongest evaluated 8B comparison method. ReCommit achieves a favorable recovery--cost trade-off, including in comparisons with the evaluated 32B models.

Mon 28 SeptArtificial IntelligenceComputation and Language
The gist
Sometimes, computer programs that use other tools to help perform tasks don’t complete what the user wants perfectly. The authors studied how to fix these programs without starting over each time something goes wrong. They developed a method called ReCommit that smartly guides the repair process by focusing on sets of possible actions rather than redoing everything from scratch. This approach lets the program find fixes more quickly and successfully, saving time and effort when fixing errors.
Open → 2609.34879v1

MACE improves multi-agent memory use for better task success

MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

Abstract: LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs needed by subsequent agents. Our empirical studies show that grouping these dependencies into functional memory units improves their retention, while connecting units increases retrieval of the units and links jointly required by a task. The preferred combination of units also changes between instructions and checklists, even when each combination's content is fixed across formats. Updating choices from the outcomes of each combination and format pairing outperforms scoring combinations and formats separately. These findings motivate MACE, a memory-agent co-evolution framework that adapts memory organization and agent memory use through execution feedback. Its MemGoG structure represents functional units as subgraphs of related conditions, actions, and outputs, connecting them through support, conflict, and repair relations. MACE Loop selects task-relevant units and relations within a memory budget and provides each agent with instructions or checklists for its current operation. It records the selected units, presentation formats, agent outputs, and task outcomes to update unit scores and relations for retrieval and inform subsequent presentation choices. Across eight benchmarks, MACE outperforms ten baselines with an average score of 81.11%, compared with 78.97% for the strongest baseline, SAGE.

Fri 18 SeptMachine Learning
The gist
Multi-agent systems made up of large language models can plan and work together on complex tasks. The authors found that breaking down the agents’ memory into connected groups called functional units helps them remember and reuse important steps better. They made a new system called MACE that adapts how memory is organized and used based on feedback from each task. MACE helps agents choose the best parts of memory and instructions during tasks, leading to better overall task performance compared to previous methods.
Open → 2609.21533v1