Papers for

ai software engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Self-improving agents hit limits unless editing tools also evolve

Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding

Abstract: An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, so frozen weights buy an eventual ceiling but no stationarity along the way. The criterion also separates three regimes usually merged: search within a fixed class, test-time training that raises the ceiling itself, and scaffold rewriting between them. Audit the scaffold, not the checkpoint. The same ceiling binds sideways. Best-of-$k$ orchestration realizes the best worker's ceiling exactly: width buys rate, not budget. Re-consulting a fixed pool has a horizon computable in advance, decided by the pool alone, and the one arrangement that would beat it, a weighted vote, needs diversity real workers lack: on 30 same-family workers the failure overlap sits at its maximum, and a majority fails 23/55 (42%) of tasks. We obtain the criterion by reading refinement as gradient boosting on the residual error between draft and target, a patch or git diff, and then measuring where that reading breaks: patches compose instead of standing beside each other to be voted on, and failures overlap. What we measure is saturation. Per-round improvement decays toward zero on SWE-bench, and churn decays geometrically across 401 production sessions, a shape shared with a pre-AI human baseline that establishes the regime without identifying its cause. Both breaks are engineering choices rather than laws about code, so together they specify a harness worth building.

Mon 28 SeptMachine LearningArtificial IntelligenceSoftware Engineering
The gist
This paper shows that when a software agent tries to improve itself by changing its own parts, it only makes small improvements unless it can also change the tools and structures it relies on. Just freezing parts of the agent doesn’t stop progress entirely, but it limits how much improvement can happen. The researchers separate three types of self-improvement: searching within fixed options, improving at run-time, and changing the tools or scaffolding around the agent. Their findings explain why improvement slows down and point to building better scaffolding as key to ongoing gains.
Open → 2609.34924v1

Large language models improve themselves by learning from their own reasoning

SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought

Abstract: Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations and constructing sophisticated environments often require substantial human effort and computation, limiting the scalability of OPSD. While a few recent studies have explored self-improvement without external PI, the resulting gains remain limited. In this work, we explore whether LLMs can achieve comparable self-improvement without external PI. Our key observation is that a single LLM can support multiple reasoning modes, such as deep-thinking and non-thinking modes, with deep thinking generating additional information during reasoning. Based on this observation, we propose Self-Evolving Online Policy Distillation (SeOPD), which enables LLMs to distill and internalize information generated by their own chain of thought (CoT). Specifically, it (1) generates CoT with the deep-thinking mode, (2) produces responses with the non-thinking mode, and (3) uses the generated CoT as PI to provide token-level supervision for the non-thinking response, allowing new information inferred during reasoning to guide the non-thinking mode and be internalized into the shared model parameters, thereby improving both non-thinking and deep-thinking capabilities. Extensive experiments across LLMs and tasks demonstrate the effectiveness of SeOPD.

Sun 27 SeptArtificial IntelligenceComputation and Language
The gist
Improving large language models usually requires outside help like expert annotations or feedback, which is expensive and slow. The authors found that a single language model can use two ways of thinking: one that thinks deeply to generate detailed reasoning steps and another that gives quick answers. Their method teaches the model to learn from its own detailed reasoning steps to make better quick answers without needing external information. This approach helps the model get better on its own over time.
Open → 2609.33181v1

Multimodal agent improves retrieval by folding context and images

MM-ContextFold: Context Folding for Multimodal Agentic Retrieval

Abstract: Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.

Sat 19 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceInformation Retrieval
The gist
Searching for information using AI agents gets tricky when they have to handle both pictures and text, as keeping all that data in one place becomes too large and confusing. The authors studied many cases and found that keeping raw images after turning them into text doesn’t help and can even cause mistakes. They created MM-ContextFold, a method that only loads images when necessary and summarizes what’s learned into text, keeping the main workspace smaller and cleaner. This approach made the agents more accurate and efficient in seven different tasks.
Open → 2609.23121v1