Papers for

ai agent developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Mcp error messages written for developers reduce success of advanced ai agents

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Abstract: Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.

Mon 28 SeptSoftware EngineeringArtificial Intelligence
The gist
Some servers that talk with AI programs send error messages meant for human programmers, telling them to run commands or change settings. The authors found that many AI agents that can only use tools get stuck because these instructions don't help them retry or fix the problem. By changing messages to name specific tools or removing tricky instructions, the authors saw these agents recover much better. This means writing error messages for AI agents needs different wording than for humans.
Open → 2609.35381v1

FlowState improves long task memory and lowers costs for AI agents

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Abstract: Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.

Mon 28 SeptArtificial Intelligence
The gist
Long tasks need AI agents to remember earlier parts of their work, but keeping all past information is expensive and squeezing it loses important details. The authors introduce FlowState, which stores and updates a special memory that keeps track of important facts and evidence separately. It lets the AI bring up old information only when needed and update decisions as new information arrives. This method improved success rates and cut down the token use by about 40% compared to remembering everything all at once.
Open → 2609.34565v1