Papers for

software tool developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Context evolution improves long task solving in AI agents

Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses

Abstract: Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.

Mon 28 SeptArtificial Intelligence
The gist
Large language model (LLM) agents often struggle with long tasks because the information they rely on grows too big and messy over time. The authors developed ContextEvo, a system that learns how to manage what information the AI focuses on by reviewing past task attempts and fixing mistakes related to context. This helps the AI keep only the most useful information when making decisions. Tests showed that ContextEvo helps AI agents do better on long, complex tasks than some earlier systems.
Open → 2609.34649v1

Context projection matches full results while cutting memory use

Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

Abstract: Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with the runner suppressing the companion whenever the first arm fails to complete. Restoring all 27 boundary runs bounds projected-minus-full success between $-$9 and +1 tasks. One omitted, selector-chosen projected continuation successfully retrieves archive text yet exhausts twelve requests; its full counterpart answers in three. The eleven jointly correct pairs form a fully observed success stratum within this recorded frame: projection reduces aggregate logical tokens by 25%, while increasing the median pair's tokens by 29% and total suffix requests from 35 to 55. Separating fitting from evaluation changes the selector's apparent tie: outside its four fitting pairs, it incurs one extra failure and 8.6% more logical tokens over thirteen comparable runs. This methodological case study connects stopping rules, known bounded failures, unexecuted companions, and resource aggregation. Its findings concern the recorded campaign, rather than population noninferiority or superiority over unrestricted Pi. Evaluations should retain every intervention boundary, execute both allocated arms independently of the first arm's completion, and report completion alongside interaction and token expenditure.

Fri 25 SeptArtificial Intelligence
The gist
The paper studies a technique called context projection that shrinks old tool outputs into small, easy-to-retrieve pieces. The authors compare this method to using full original data in a coding task environment. They find that both ways succeed equally often, but the projected method uses less overall memory while sometimes needing more interaction. Stopping rules and replaying both methods independently matter for fair evaluation. This work focuses on a specific recorded experiment rather than proving general superiority.
Open → 2609.31381v1

Large language models fix bugs by rewriting more code than humans do

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

Abstract: Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions ($\sim$ 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.

Thu 24 SeptComputation and LanguageSoftware Engineering
The gist
Fixing bugs in computer programs can be done by either patching the broken code or writing new solutions. The authors studied how well large language models (LLMs) fix buggy code compared to humans. They found that LLMs often change more lines than needed and sometimes rewrite entire solutions rather than just fixing bugs. LLMs also perform better when allowed to create new code from scratch instead of trying to patch buggy code. This insight matters for building programming tools that help developers fix code incrementally.
Open → 2609.29410v1

FeatLens cuts code graph size to speed up repo level code generation

FeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation

Abstract: Recent code generation research has moved from isolated function completion toward repository-level generation in existing codebases. To implement a target function correctly, an LLM must identify reusable repository dependencies such as existing functions, APIs, and cross-file definitions. Existing retrieval methods provide such context through code similarity search, persistent whole-repository graphs, or LLM-driven graph exploration, but often incur high graph construction, reasoning, and token costs. Feature-oriented methods offer a natural view of software functionality, yet they mainly support requirement decomposition, planning, or feature editing rather than code dependency retrieval. This paper presents \textbf{FeatLens}, a feature-guided dynamic code graph construction and retrieval approach for repository-level code generation. FeatLens builds a feature index that links natural-language feature descriptions to function-level code entities. Given a generation task, it dynamically constructs a task-specific seed graph from the feature index and applies semantic-structural graph reasoning with personalized PageRank to select a compact reasoning graph. This design replaces persistent whole-repository graph maintenance and LLM exploration with deterministic and lightweight dependency retrieval. Experiments on DevEval and EvoCodeBench show that FeatLens achieves the best DR@15 among sparse, dense, and graph-based baselines (0.501 and 0.460). On DevEval generation, it obtains the highest DIR@1, reaching 52.91\% with DeepSeek-V3.2 and 53.58\% with GPT-5-mini, while maintaining competitive Pass@1 and producing shorter code. Compared with the strongest graph-based baseline, FeatLens reduces graph nodes by 61.0\%, edges by 86.2\%, and total token overhead by 45.9\%, with no LLM tokens used during retrieval.

Tue 22 SeptSoftware EngineeringArtificial Intelligence
The gist
Generating new code inside a large software project requires understanding which existing code parts can be reused. The authors present FeatLens, a method that uses natural language descriptions of features to quickly find the most relevant code pieces and their connections. This avoids building huge, slow-to-handle graphs of the entire project and helps code generation models focus on a smaller set of important code units. Their experiments show that FeatLens improves the retrieval of needed code dependencies and reduces overhead compared to previous approaches.
Open → 2609.26480v1

Machine learning prototypes require major changes to reach production

"It Comes in Notebooks": Changes and Challenges when Operationalizing ML Prototypes

Abstract: Machine learning practitioners commonly prototype models in computational notebooks before transitioning them to automated production systems. Despite its prevalence, the concrete engineering work involved in this transition and the software quality concerns that motivate it remain insufficiently characterized. We report on a qualitative study based on semi-structured interviews with 13 ML practitioners from industry and academia. Using reflexive thematic analysis, we identify 23 engineering changes organized into five themes: code restructuring, data pipeline development, testing & validation, pipeline automation, and monitoring & observability. We also identify 20 software quality attributes across the ML development lifecycle and map them to the engineering changes. A recurring pattern in our findings is that computational notebooks externalize oversight to the human practitioner, and defer costs that become obligatory at operationalization time. Operationalization constitutes the repayment of this technical debt accumulated during prototyping, which we refer to as oversight debt. Practitioners do not merely restructure notebook code, but repay this debt by constructing automated substitutes for the interactive oversight that notebooks provide. We further present seven quality trade-offs showing that these tensions are properties of the notebook-to-production transition, rather than symptoms of poor engineering practice. Our findings structure operationalization effort, establish empirical links between engineering changes and software quality concerns, and provide implications for practitioners, tool designers, and researchers working on ML-enabled software systems.

Sat 19 SeptSoftware Engineering
The gist
Working on machine learning models often starts in interactive notebooks where developers explore ideas. The authors found that changing these notebooks into reliable, automated systems involves many essential engineering tasks that address quality and monitoring. These tasks repay the 'oversight debt' built up when notebooks let people do much of the checking manually. The study identifies common engineering changes and how these relate to software quality in machine learning projects.
Open → 2609.22903v1

Software supply chain security needs unified risk measurement methods

Measuring the Security of the Evolving Software Supply Chain: a Research Agenda

Abstract: Software supply chain security has become increasingly critical due to the widespread reliance on third-party dependencies and the growing attack surface of modern software ecosystems. However, existing quantitative, measurement-based analysis and vulnerability management approaches remain largely fragmented and ecosystem-specific, limiting their ability to provide comparable risk assessments across environments. This paper presents a structured research plan, starting with a Systematization of Knowledge (SoK) to synthesize the current state of research and identify key gaps, highlighting the limitations in dependency modeling and vulnerability propagation analysis, particularly in the treatment of transitive dependencies and their real-world exploitability. Based on these insights, we argue for a unified measurement perspective capable of consistently representing and analyzing the cross-ecosystem dependency structure. We further identify emerging challenges introduced by AI-assisted software development, where coding LLMs are likely to contribute to new dependency patterns that are not captured by traditional Software Composition Analysis (SCA) tools. These shifts motivate a rethink of dependency modeling to account for evolving software-generation practices and their long-term structural impact on software security.

Tue 8 SeptCryptography and SecuritySoftware Engineering
The gist
Modern software relies heavily on code from many other sources, creating a complex web of dependencies that can have hidden security problems. The authors find that current tools to measure risks in these software supply chains don't work well across different software ecosystems and often miss how vulnerabilities spread, especially in indirect dependencies. They suggest a unified way to measure and understand these connections more consistently. Also, they point out that new ways to write code using AI tools may introduce new types of dependencies that existing security checks don't catch.
Open → 2609.08810v1