Papers for

customer support automation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Failure-transparent agents reduce false success claims in AI tools

Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

Abstract: Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.

Mon 28 SeptArtificial Intelligence
The gist
Sometimes AI agents use tools that fail, and then the agents incorrectly say they succeeded without proper proof. The authors created a special test called Failure-Transparent Agents (FTA) to focus on checking whether agents honestly report failure after tools fail. They tried different models and methods, showing that adding clear instructions and requiring structured evidence drastically reduces false success claims and made responses more helpful. This work helps improve trustworthiness in AI that uses other tools.
Open → 2609.35732v1

Model uses confident answers to stop reading early and save time

The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading

Abstract: Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model whether it has read enough. We introduce Answer-Convergence Stopping (ACS), a training-free stopping rule that measures rather than asks. After each chunk, it probes the frozen model's current answer state and stops when that state is both confident and stable. The rule requires only output-side generation and token log probabilities, has no trained components, and uses one shared configuration across models and benchmarks. Because a stopping policy can save computation simply by stopping too early, we evaluate the stopping decision itself using evidence position where available. On the full LongBench-v2 with two frontier models, ACS is the only stopping policy that matches or exceeds full-reading accuracy. Furthermore, across 250 S-NIAH questions, the premature stopping rate for ACS across five models from two families ranges from 0% to 12%, compared to 8.4% to 45.6% for the verbalized gate. Taken together, ACS reveals that by properly utilizing the output signals of frozen models, we can achieve favorable behaviors like adaptive stopping without the need for additional training.

Mon 28 SeptComputation and LanguageMachine Learning
The gist
Processing long texts can take lots of computer time if a language model always reads everything fully. The authors create a simple way for a model to decide when it has enough information by checking how confident and stable its answers are, without any extra training. This method, called Answer-Convergence Stopping, saves time by stopping early without losing accuracy. They tested it on tough reading tasks and found it works better than other methods that require extra learning.
Open → 2609.34590v1

Memory system for AI agents improves itself before errors happen

Remember Before You're Asked: MemDream for Self-Probing Memory Evolution

Abstract: Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.

Mon 28 SeptArtificial Intelligence
The gist
Language models that act as agents need memory to keep track of information over long tasks, but usually they only fix memory mistakes after they cause problems. The authors propose MemDream, a method where the system pretends to test its own memory when it's not actively working, finding and fixing issues early. This approach uses three roles working together to analyze and improve memory proactively, resulting in better answers and performance on tests compared to older methods. MemDream helps AI agents remember better before real mistakes occur.
Open → 2609.34545v1

Fixed language model judges make version-dependent mistakes evaluating agents

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

Abstract: Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned agent-task cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare improvements that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge confidently reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and reference outcome, while a task-solvability prediction from AgentRewardBench reverses sign in SWE-bench. Transporting old-version calibration raises mean absolute comparison error on SWE-bench from 3.8 to 19.5 percentage points; 24.6% of ratio-bootstrap draws are undefined near the correction boundary. A tuned paired audit narrows a classical interval by only about 5% at 80 labeled tasks. A randomized three-arm test does not support the predicted increase in false acceptance from showing the agent's final report (all three Holm-adjusted p-values = 1.0). These results favor explicit reference standards and paired audits of current outputs over judge-only release decisions or transported old-version calibration.

Mon 28 SeptMachine Learning
The gist
Evaluating improvements in AI agents using language-model judges that are kept the same over time can lead to errors that depend on which versions of the agents are compared. The authors studied multiple coding and customer-service AI agents and found that the judges often disagree depending on agent versions, sometimes wrongly approving inferior upgrades. They suggest that relying solely on these fixed judges or old calibration can cause significant mistakes. Instead, they recommend comparing agent outputs directly with explicit reference standards and current data audits.
Open → 2609.34198v1

Recursive skill evolution improves large language model agents performance

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

Abstract: LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.

Sun 27 SeptArtificial Intelligence
The gist
Large language model (LLM) agents can get better at tasks by reusing and improving the skills they combine to solve problems. The authors present a method called R² Flow that helps these agents improve themselves step-by-step, by learning, verifying improvements, and updating their skill libraries more reliably. This method organizes how skills are combined so the system can share learning across similar task paths and decide more accurately which skills to change. In tests involving question answering, math reasoning, decision making, and code generation, R² Flow led to better accuracy and more precise skill updates than previous approaches.
Open → 2609.33867v1

Evolution-aware memory improves long-term AI agent interactions

EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval

Abstract: Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks. However, existing memory systems struggle to utilize continuously evolving historical information, as they often rely on static memory representations and single-round retrieval strategies, failing to track factual changes or integrate distributed evidence across long-term interactions. To address these challenges, we propose \textsc{EMIR}$^{2}$, an \textbf{E}volution-Aware \textbf{M}emory framework with \textbf{I}ntent-Guided Multi-\textbf{R}ound \textbf{R}etrieval, enabling LLM agents to maintain evolving historical knowledge and adaptively retrieve relevant evidence. Specifically, \textsc{EMIR}$^{2}$ constructs a State-Evolving Memory Graph (SEMG) that represents long-term memory as evolving knowledge states supported by temporal event trajectories and evidential associations. By maintaining semantic states through evidence-based updates, SEMG preserves historical evolution and enables evidence tracing under complex and conflicting scenarios. Building upon this, we introduce an intent-guided multi-round retrieval mechanism that iteratively identifies missing evidence and expands retrieval based on accumulated information. Experiments on LoCoMo and MemConflict demonstrate that \textsc{EMIR}$^{2}$ improves long-term memory utilization, dynamic and static conflict handling, and complex retrieval performance, achieving relative improvements of more than 12\% in certain categories. These results highlight the effectiveness of jointly modeling memory evolution and adaptive evidence acquisition for long-term agent interactions.

Sat 26 SeptArtificial Intelligence
The gist
When AI agents try to remember things for a long time, their memory can get outdated or miss important details that change over time. The authors propose a new memory system called EMIR2 that keeps track of how information evolves and uses multiple rounds to find the right details. This system helps AI handle changes in facts and gather complex evidence better than before. Tests show it improves AI memory use and decision-making by over 12% in some cases.
Open → 2609.32584v1

Multi-agent systems manage changing tasks with issue tracking

RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems

Abstract: LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution. We refer to such problems as progressively specified tasks. To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence. We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management. RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving. Across ProgSpec and five existing benchmarks, RepoMAS achieves the best performance. Further analyses show that its issue-driven revision and repository maintenance mechanisms consistently contribute to performance. These results highlight the importance of allowing MASs to revise not only how a task is solved, but also revise their explicit representation of task requirements during execution.

Sat 26 SeptArtificial Intelligence
The gist
Tasks often change or get clearer while people are still working on them, but many computer systems expect all task details beforehand. The authors studied tasks that are only partly known at the start and become clearer during work. They created a new test that checks if a system can handle these changing needs well. Their system, RepoMAS, keeps track of new problems and changing requirements like a project manager would, updating its plan as it works. This approach performed better than others on several tests, showing that computer agents benefit from tracking and revising their understanding of tasks as they go.
Open → 2609.32490v1

Enterprise agents struggle with hidden facts in company data questions

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Abstract: In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage. For each generated company, code fills each template and computes an exact answer without a language model. We evaluate 12 agents. Each pairs a model with an agent program, which connects it to the company's systems. The best agent answers 18 of its 24 attempts, three per question, correctly. Four of the six models answer at most 6 of 24 with any program. The hardest questions require picking one of several similar records, such as which of three renewal offers a customer signed. All agents together answered two such questions correctly in only 1 of 84 attempts.

Thu 24 SeptSoftware EngineeringArtificial Intelligence
The gist
Some questions about company data have facts that aren't directly stated, making them hard for AI agents to answer. The authors created a test where agents must find answers relying on hidden clues, not just stated facts. They found most AI models handled easy questions well but struggled greatly with these hidden fact questions. This shows current AI tools have trouble understanding subtle or implied information in business settings.
Open → 2609.30055v1

Framework improves math tutoring by learning from past conversations

REAT: A Reflective Experience-Augmented Tutoring Framework for Multi-turn Mathematical Instruction

Abstract: Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.

Thu 24 SeptMultiagent Systems
The gist
Solving math problems well doesn’t always mean a computer can teach math effectively. The authors designed a system called REAT that learns from previous tutoring chats and uses what it has learned to help students better during live sessions. It watches past conversations to find good teaching patterns and then uses those at the right moments, adapting to how each student is doing. Tests show that this method helps more in tough teaching moments than just giving the computer fixed instructions or retraining it.
Open → 2609.29804v1

Typed classifier judges answers cheaper and faster than llms with similar mistakes

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

Abstract: We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev's accuracy differs significantly from an LLM judge's in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed over the nine panels, the LLM judges, called once per criterion, cost 29 to 325 times as much as Jev and took 30 to 220 times as long. On graded criteria all four judges agree more with one another than with the labels and mostly assign lower levels than the raters. One of several observational accounts is that raters followed scale conventions our criterion texts omit. Jev's confidence ranks its own errors on most panels, which should make a cheap classifier the ideal first stage of a cascade that defers its uncertain verdicts to an LLM judge. Correlated errors undo that advantage. The LLM judges repeat nearly all of Jev's most confident errors, so a cascade replayed on the recorded verdicts lowers cost but gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 even with oracle thresholds.

Thu 24 SeptComputation and Language
The gist
The authors investigate whether a typed classifier called Jev can be used instead of large language models (LLMs) to judge answers based on rubrics. Jev is much cheaper and faster than LLM judges and performs similarly in accuracy, especially on simple yes/no criteria. However, both Jev and LLM judges tend to make similar mistakes, especially on more graded, complex criteria. This means using Jev first and then asking an LLM only on uncertain answers saves cost and time but improves accuracy only a little.
Open → 2609.29769v1

SkillAA improves AI skill updating with precise graph-based editing

SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

Abstract: External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.

Thu 17 SeptArtificial Intelligence
The gist
AI systems often use fixed skills to perform tasks, but fixing errors in those skills can be tricky and inefficient. The researchers created SkillAA, a method that uses a detailed graph to represent skills and how they connect, allowing the system to locate the exact part that caused a failure and update only that part. This makes AI skill updates more accurate and reliable. Testing with a powerful language model showed SkillAA achieves high success rates on several question-answering tasks.
Open → 2609.20455v1

Multi-turn agents refined to cut cost and boost accuracy

Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

Abstract: Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost. We propose viewing each trajectory as a \emph{round-level dependency DAG} that exposes which rounds are globally load-bearing for the final answer, and fine-tune agents on trajectories refined through this DAG. Given an LLM-annotated DAG, these edits are deterministic and interpretable, with optional rephrasing. Models trained on these refined trajectories consistently outperform those trained on the original trajectories at lower inference cost. Specifically, across four multi-modal QA benchmarks, our refinements improve downstream accuracy by up to $1.7$\,pp over vanilla SFT (and $5.7$\,pp over an LLM-deletion baseline) while reducing per-sample inference messages by up to approximately $40\%$ and inference tokens by up to approximately $48\%$, translating to substantial savings in compute and serving cost. Code is available.

Wed 16 SeptComputation and Language
The gist
When computers answer questions in multiple steps, some steps are unnecessary and waste time and resources. The authors suggest looking at these steps like a map that shows which ones really matter for the final answer. By teaching computers only with the important steps, they become better and faster at answering questions. This method was tested on several question-answering tasks and improved accuracy while reducing the time and computing power needed.
Open → 2609.18417v1

Speech language models improve reasoning accuracy with real time self correction

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

Abstract: Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance reasoning capabilities without inducing prohibitive delays, an inherent accuracy-latency trade-off persists. In this paper, we investigate whether a streaming SpeechLLM can dynamically revise its reasoning traces on the fly. We introduce RetroThinker, a multi-stage post-training framework that equips the Moshi model to self-verify and forward-correct CoT steps during inference. RetroThinker combines supervised fine-tuning (SFT) on curated retrospective thinking data with length-based direct preference optimization (DPO) to optimize retrospective during early reasoning (i.e., reasoning concurrently while the user speaks). Evaluated on the GSM8K benchmark, RetroThinker significantly improves the accuracy-latency trade-off over non-retrospective baselines, achieving an 11% absolute accuracy gain at a comparable latency.

Thu 10 SeptArtificial IntelligenceComputation and Language
The gist
Speech language models can understand spoken words faster and keep voice details better than converting speech to text first. But they are not as good as text-based models at solving tricky problems quickly. The authors created RetroThinker, which lets a speech model check and fix its own thinking steps while listening. This approach helps the model be more accurate without slowing it down much. Tests showed RetroThinker improved problem-solving scores by 11% with similar speed.
Open → 2609.11864v1

Weaker AI models improve task success with expert-guided corrections

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Abstract: Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harness and model weights shape behavior, we ask how harness evolution and lightweight fine-tuning should be combined. Across seven enterprise agent tasks, we first evolve a harness with the weaker model, then find that a stronger expert often uses it more effectively, suggesting expert supervision could close the remaining gap. However, training the weaker model on the expert's complete trajectories under the evolved harness backfires: performance regresses on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, even though the same procedure helps under the unevolved harness. Our analysis shows that imitation transfers knowledge and increases scaffold usage, but disrupts model-harness fit: the weaker model adopts the expert's planning strategy without the competence to execute it and no longer matches the harness evolved around its native planning style. We therefore develop an on-policy expert-correction pipeline, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn. This preserves the model's planning style and combines the gains of harness evolution and model adaptation. Our results identify and resolve a source of contention between harness and weight updates, yielding a compatibility-preserving recipe for economical co-evolution on domain-specific enterprise tasks.

Tue 8 SeptArtificial Intelligence
The gist
When smaller AI models try to do complex tasks, they rely on surrounding tools and instructions called harnesses to help them succeed. The authors found that just copying how a stronger AI model plans often makes the weaker model perform worse because the weaker model cannot do everything the stronger one does. To fix this, they created a method where the weaker model runs its own plan but gets help only on parts where it struggles, using corrections from the stronger model. This way, smaller models can learn better and work well with their own harness setup.
Open → 2609.09134v1

Memory cleaning improves task success for long running AI conversations

MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents

Abstract: Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences and task knowledge across extended interactions. Conventional retrieval mechanisms optimize semantic compatibility rather than downstream utility, frequently introducing outdated, misleading, or conflicting evidence into the active context. We present MeClear, a task conditioned memory clearance framework that identifies memories featuring negative downstream utility through cooperative attribution and selectively suppresses them from agent execution. MeClear combines Leave One Out screening with sampled cooperative Shapley attribution to distribute utility across interacting evidence, effectively resolving redundant conflict masking where single removal evaluations fail. Utilizing attribution rankings, MeClear executes a query scoped minimal clearance strategy over a nested filtration, verifying task recovery on the cleared context without permanently altering the persistent memory bank. Comprehensive experimental evaluations across ten long dialogue memory pools demonstrate that MeClear achieves a target recall of 85.9% and an overall task recovery rate of 82.3%, representing a 25.5 percentage point improvement over Leave One Out (LOO) baselines.

Tue 8 SeptArtificial IntelligenceSoftware Engineering
The gist
When AI systems like chatbots remember too much over a long time, they sometimes use outdated or conflicting information that confuses their responses. The authors introduce MeClear, a new method that carefully checks which remembered information actually helps the current task and removes the parts that hurt the AI’s performance. By doing this in a smart way, MeClear improves how well AI agents complete complex tasks over long chats. Their tests show it works better than existing methods that remove bad memories one by one.
Open → 2609.09115v1

AgentGrad improves multi agent system prompts with guided fixes

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

Abstract: Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language feedback have emerged as a leading paradigm. In this paper, we identify limitations in two stages of existing textual gradient approaches: gradient extraction and gradient aggregation. In gradient extraction, previous works select a target prompt without verifying whether modifying it resolves the failure, and derive gradients without agent-level supervision over the corresponding agent's intermediate output. In gradient aggregation, individual gradients are randomly grouped and concatenated, often mixing unrelated failure modes and producing prompts that fail to generalize. To address these limitations, we propose \textbf{AgentGrad}, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction. For each failure, sequential intervention modifies the behavior of one agent at a time to identify the target agent whose modification resolves the failure. The modified output of the target agent then serves as agent-level supervision for extracting a fine-grained gradient. Semantic textual gradient abstraction clusters semantically similar gradients to prevent mixing unrelated failure modes, and abstracts each cluster into a generalized gradient that captures the shared corrective pattern. Experimental results show that AgentGrad achieves state-of-the-art performance across five MAS benchmarks and reduces wall-clock optimization time by $2.5\times$ on average compared to the next-fastest baseline.

Tue 8 SeptArtificial Intelligence
The gist
Large language models work better when many specialized agents collaborate, but their success depends on how well each agent is told what to do. The authors found problems in existing ways to improve agent instructions because they guess which agent caused errors and mix fixes that don’t relate. They created AgentGrad, which tests agents one by one to find the one causing the problem and groups similar fixes for clearer improvements. This method works better and speeds up making agents smarter across several tests.
Open → 2609.08572v1