Papers for

devops teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Compiler can access code part but coding agent cannot enforce it

The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent

Abstract: The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.

Mon 28 SeptCryptography and SecurityArtificial IntelligenceSoftware Engineering
The gist
Sometimes, parts of computer code need to be hidden from helpers that write or modify code, but still accessible to tools that prepare the program for running. The authors looked at different ways to control who can read these hidden parts and found that common methods like containers, permissions, and sandboxes cannot tell who is reading the code. This means it's hard to protect some intellectual property in code from being seen by automated coding assistants, even though it's needed by compilers.
Open → 2609.35557v1

Planarian improves AI agent state management for better task recovery

Planarian: Managing Agent State with Statepoints

Abstract: LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from erroneous ones. Doing so safely requires coordinated actions, yet current agent harnesses lack unified abstractions and mechanisms for managing local and remote state consistently and efficiently. We describe Planarian, an agent runtime with state management that enables agents and users to recover from erroneous actions and explore alternative executions over consistent local and remote environment state. Planarian introduces the abstraction of agent statepoints, which are consistent, restorable point-in-time versions of the environment state. Planarian exposes three state-management primitives to agents and users: (i) snapshot creates a new statepoint spanning local and remote state without requiring external services to support checkpoints: it relies on efficient incremental process and file system snapshotting to capture local sandboxed state, and transparently records compensating actions to undo remote state changes; (ii) rollback restores the environment to a previous statepoint by reverting to a prior local checkpoint and replaying compensating actions for remote state changes; and (iii) fork creates multiple isolated branches from a statepoint, enabling the agent to explore alternatives in parallel. We show that Planarian enables agents to undo mistakes and explore alternatives in parallel, improving task quality by up to 15x, and allows users to recover from erroneous actions with only 3% overhead.

Mon 28 SeptOperating SystemsArtificial IntelligenceCryptography and Security
The gist
LLM agents work by making changes across files, tools, and online services, which can get messy and hard to fix if errors happen. The authors created Planarian, a system that lets these agents save snapshots of their state, undo mistakes, and try out different options easily. Planarian keeps local and remote parts of the agent’s environment in sync and safe, which helps agents and people recover from problems faster. This approach improved task success by up to 15 times while adding little overhead.
Open → 2609.35366v1

Coding agents often duplicate code despite functional success

Do Coding Agents Reuse Existing Code or Reinvent the Wheel?

Abstract: Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8\% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.

Mon 28 SeptSoftware EngineeringArtificial IntelligenceComputation and Language
The gist
Coding agents that write software sometimes copy the same code multiple times instead of reusing existing code. The authors built a benchmark called RepoReuse to check if these agents reuse code or reinvent it during multiple steps of software development. They found that agents frequently leave duplicated code even when their code passes all tests, which could cause extra maintenance work. This shows that just checking if code works well is not enough to evaluate coding agents.
Open → 2609.35357v1

Graph guided method improves software environment setup success

Graph-Guided Repository Environment Construction

Abstract: Coding agents now increasingly rely on execution to validate their solutions, making the construction of reliable execution environments a critical enabling capability. However, repository environment construction is challenging because execution requirements are fragmented across repository artifacts and may only become apparent during execution. Existing agent-based approaches address this problem through iterative interaction, but information about the current construction state, including discovered requirements, satisfied and unresolved prerequisites, and their dependencies, can remain distributed across the interaction history. We present Graph2Env, an agent-based approach centered on DepGraph, a typed dependency graph that explicitly represents the environment requirements needed for repository execution, their dependency relations, and their states. Graph2Env uses DepGraph to guide environment construction and continuously refines it with execution feedback, while persisting successful repairs into a replayable construction procedure. The resulting artifacts are then applied in a fresh environment to verify that the constructed environment can be reproduced. We evaluate Graph2Env on a benchmark of 200 Python repositories drawn from RATBench and EnvBench, against a static dependency-inference baseline (pipreqs), three specialized environment-construction systems (Repo2Run, RAT, and SetupX), and two general-purpose coding agents (SWE-agent and Claude Code). Graph2Env achieves an 81.0% Environment Build Success Rate (EBSR) and a 59.3% Environment Setup Success Rate (ESSR), outperforming the strongest baseline by 9.5 and 9.0 percentage points, respectively.

Sun 27 SeptSoftware EngineeringArtificial Intelligence
The gist
Getting software project environments ready to run code properly is hard because requirements are spread across different files and may only show up when testing the code. The authors created Graph2Env, a tool that maps out all the needed parts and their connections in a clear graph to help build these environments step-by-step. By learning from errors during execution and keeping track of what works, Graph2Env builds setups that can be repeated reliably. When tested on many Python projects, it did better than existing methods at preparing working environments.
Open → 2609.33429v1

Sandbox security flaws in ai orchestration expose high risk exploits

Weird Machine Compositors: Exploiting AI Orchestration at the Expression Layer

Abstract: Orchestration platforms secure user-provided expressions through enumerate and block sandboxing: AST rewriting, runtime property blocklists, template sandbox environments. We demonstrate that these sandboxes are weird machines whose instruction set is the underlying language specification, and that the enumerate and block approach is unfixable, following the same trajectory that led to the deprecation of past sandboxing technologies such as Java's SecurityManager and vm2. We validate this claim through three rounds of escalating bypasses against n8n's expression sandbox (three CVEs, two CVSS 9.4, one unauthenticated), and frame these findings within a broader pattern of sandbox failures across the orchestration products category. We identify a trust laundering pattern where orchestration pipelines and applications move attacker controlled input from untrusted to fully credentialed through transformations that strip taint at each level. AI-assisted enumeration accelerates the discovery of these coverage gaps, compressing the timeline between a sandbox's deployment and its compromise. We provide an AST coverage analysis methodology, an accompanying open-source tool, and a defensive playbook that includes policy inversion (allowlist over blocklist) as a structural mitigation.

Sun 27 SeptCryptography and Security
The gist
The paper shows that common security measures used to lock down user expressions in AI orchestration platforms are not effective. The authors demonstrate that these sandboxes behave like 'weird machines' that can be manipulated to bypass restrictions. They found multiple serious vulnerabilities in a popular platform called n8n and explain how attackers can move from untrusted inputs to fully trusted access by tricking these systems. They also show that AI helps attackers find these weaknesses faster. The authors suggest a better defense approach that focuses on allowing only known safe operations instead of blocking certain unsafe ones.
Open → 2609.33413v1

AI code fixes overload review and build systems in big projects

Orchestrating AI-Assisted Code Remediation: Socio-Technical Bottlenecks in a Large Industrial Repository

Abstract: Background: Code degradation in large, long-lived codebases is costly to remediate through manual refactoring and opportunistic clean-ups. LLM-based coding assistants can perform mechanical remediation at scale, but their impact on industrial workflows is underexplored. Objective: We investigate how massive AI-assisted code remediation affects build-on-commit continuous integration (CI), code review, and team coordination in a large industrial repository, and which socio-technical bottlenecks constrain such remediation when source editing becomes cheap through AI assistance. Method: We report on a 15-day exploratory single-case field study in which an experienced developer used a command-line AI coding buddy to remediate widespread issues in a closed-source industrial C++ repository. We triangulate Gerrit metadata with a developer diary and team chat, analyzed through descriptive statistics and qualitative coding. Results: AI-assisted remediation rapidly generated hundreds of commits touching thousands of lines, saturating CI and reviewer attention. Naïve per-file commits overloaded build-on-commit CI; Switching to directory-based batching and capping the number of files per change restored throughput, but still required explicit review solicitation, negotiation of acceptable commit granularity, and iterative follow-up to resolve build and static-analysis failures. Conclusion: When mechanical editing is cheap, CI capacity, review effort, and change orchestration become primary bottlenecks. Sustainable AI-assisted remediation in very large repositories requires deliberate control of commit, review, and CI batch granularity and treating semantic change sets, such as ``fix all instances of warning X'', as first-class units of work that can be sliced differently for developers, reviewers, and CI.

Thu 24 SeptSoftware Engineering
The gist
Fixing lots of code problems by hand is slow and costly. The authors studied how using AI tools to quickly fix code issues in a large industrial project caused new challenges. These fixes created many code changes that overwhelmed automated building and human review systems. The study found that managing how and when code changes are grouped and reviewed is very important when AI makes editing easy. This helps keep the repair process smooth and sustainable in big software projects.
Open → 2609.29172v1

Algebraic architecture theory helps understand software changes reliably

Foundations of Algebraic Architecture Theory: A Rising Sea of Geometry, Transport, Comparison, and Reconstruction

Abstract: AI-generated software changes make it increasingly important to determine what a change preserves, where local consistency fails to extend globally, and which alternatives remain. We develop the foundations of Algebraic Architecture Theory (AAT) from Atoms, typed primitive facts, and Laws, equations that objects must satisfy. A reading specifies what counts as structure and which operations and laws to preserve. The main reconstruction theorem identifies the category of full geometries and all their structure-preserving morphisms with an independently defined category of local models, up to equivalence. Objects are recovered up to isomorphism and morphisms between fixed endpoints uniquely. The theory addresses gluing, diagnosis, transport, classification of changes, and reconstruction. From finite Atom families we construct cores closed under operations and geometries with sites and coefficients. We give conditions under which a Cech obstruction detects the existence of a global state and, through comparison with repair semantics, a global repair. We compare diagnoses and give a finite criterion for uniform invariance given computable finite data. Transport along exact changes has a universal property and commutes with base change on exact pointed pullback squares. Comparisons of routes generated from the same square, finite comparison diagram, and geometry factor into an invertible comparison and an idempotent normalization. We characterize when observations determine comparison preservation and classify compatible lifts. Encodings of lens and protocol semantics preserve and reflect laws and recover semantics-preserving morphisms. Applications classify and count operation-preserving changes and extend morphisms uniquely from finite tables. Corresponding Lean declarations are listed in the appendix.

Wed 23 SeptSoftware EngineeringProgramming Languages
The gist
Software changes made by AI need to be examined carefully to see what parts stay the same and where problems might appear. The authors develop Algebraic Architecture Theory (AAT) to capture the structure and rules that software must follow, helping to reconstruct and compare different versions. Their main result shows how local pieces relate to the whole, making it easier to detect if something breaks globally or can be fixed. This theory also helps classify changes, diagnose issues, and transport information accurately between versions.
Open → 2609.27638v1

Benchmark evaluates AI agents configuring software deployment environments

FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

Abstract: Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes. Agents submit declarative artifacts that are collected, rebuilt, and redeployed in a pristine environment. Four gated binary check layers measure build, readiness, behavior, and conformance to the deployment specification, using programmatic checks without an LLM judge. A four-arm release gate requires a resolving reference solution and rejects tasks solved by do-nothing, specification-transcription, or generic-stub submissions. The released check annotations expose the link between 2,145 checks and their specifications, including seven documented gaps. Three additional adversarial strategies test shortcuts in the grading signals; none resolves any of the 135 tasks they cover, while a vacuous health probe passes readiness and exposes the need for downstream checks. On the 136-task evaluation grid, seven language models from four providers use the same four-tool scaffold and resolve 52.9-75.0 percent of tasks. The three zero-intelligence floors resolve none and reach a mean Deployment Score of at most 0.44. Readiness is the largest failure stage, accounting for 110 of 313 unresolved episodes. Mean resolution rate is 30.7 percentage points higher on the repair task group than on the disjoint greenfield group, with a positive gap for every model; ten tasks resist all seven. In a 25-task case study, one practicing engineer directing Claude-Sonnet-5 resolves 92 percent against 72 percent for the autonomous baseline. FDE-Bench links deployment success and failure to artifacts that can be inspected and replayed.

Wed 23 SeptSoftware EngineeringArtificial Intelligence
The gist
Setting up software so it runs properly and stays working can be tricky. The authors created a test suite called FDE-Bench that checks how well AI systems can turn code into working software setups using technologies like Docker and Kubernetes. Their tests include building, starting, and checking if the software works as expected, without using AI to judge the results. They found AI can solve many but not all deployment tasks, especially struggling with readiness checks. The benchmark also helps engineers understand and improve deployment processes by linking successes and failures to specific configuration artifacts.
Open → 2609.27571v1

Kubernetes security misconfigurations identified and fixed with AI models

Kubernetes Misconfigurations in the Wild: Taxonomy, Evolution, and Automated Repair with Large Language Models

Abstract: Kubernetes is widely used to orchestrate cloud-native applications, yet its declarative configuration model often introduces security misconfigurations that threaten system reliability. Despite available detection tools, misconfiguration patterns and scalable remediation remain insufficiently understood. This paper presents an empirical study of Kubernetes security misconfigurations based on 2,662 developer-reported Stack Overflow issues. We derive a taxonomy of recurring security weaknesses across configuration objects and categories. We analyze severity variations and investigate how misconfigurations evolve between incubator and stable project stages. Findings show that while some operational issues decrease as projects mature, critical security misconfigurations often persist or reappear. We then evaluate Large Language Models (LLMs) for automated remediation under progressively enriched contextual conditions. Contextual grounding improves correction accuracy, with the best standalone model achieving 89.06%. To enhance structural correctness and schema compliance, we introduce Kubecurity, a schema-guided validation framework based on official Kubernetes specifications. Combining contextual LLM reasoning with deterministic schema enforcement achieves 98.50% correction accuracy while substantially reducing newly introduced misconfigurations. This work advances the understanding of Kubernetes security misconfigurations and demonstrates a hybrid approach to more reliable automated remediation.

Tue 22 SeptSoftware Engineering
The gist
Kubernetes is software that helps run applications in the cloud, but setting it up can cause security mistakes that make systems unsafe. The authors studied many real problems reported by developers online to find common security errors and see how they change over time. They tested large language models (AI that understands and writes code) to automatically fix these mistakes and created a tool to check fixes against official rules. Together, the AI and the tool corrected nearly all errors while avoiding new ones, helping make Kubernetes setups safer.
Open → 2609.27030v1

Category-aware training improves software engineering AI performance

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

Abstract: Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the training pools. Initial category-specific RL improves average training success while leaving uneven instance-level progress, motivating explicit consolidation of successful behavior and policy-adaptive task selection. Same-origin category experts alternate long-horizon Agentic-miniRL with Refresh-Repair-Expand (RRE): the updated policy refreshes instance mastery, reuses its own verified successful trajectories for Repair SFT, and reselects tasks for further RL. Label-routed multi-teacher on-policy distillation (MOPD) consolidates the experts into one deployable student, with ReLU-gated reward extrapolation keeping only each teacher's improving direction over the reference. Expert training and policy integration require no external model to provide solution trajectories or action targets. We evaluate Pooled RL and Balanced RL, expert development, and single-model integration through aggregate and per-category resolution, the minimum category lift over each joint-RL baseline, and expert-gain recovery. The final MOPD policy achieves mean resolution of 58.04% on Pro-618 and 59.00% on SWE-bench Multilingual, improving over the base model by 5.39 and 2.78 percentage points, respectively.

Sun 20 SeptSoftware EngineeringComputation and LanguageMachine Learning
The gist
Software engineering tasks can be very different from each other, and when AI agents are trained on many tasks at once, some tasks improve while others get worse. The authors created a way to train AI experts for specific categories of software tasks, then combine their knowledge into one smarter AI. This method helps balance improvements across all software tasks and avoids losing progress on any one type. Their experiments show that this approach makes the AI better at solving diverse programming problems than previous methods.
Open → 2609.23377v1

LLM program repair results need clear experiment details

Experimental Settings in LLM-Based Program Repair: A Study of Inputs, Tool Access, Feedback, and Validation

Abstract: Evaluations of automated program repair (APR) systems commonly report the benchmark, the number of repaired defects, and the tests used for final patch validation, but these items no longer fully specify the repair task presented to a system. Recent LLM-based systems differ in the information supplied before repair, the repository and testing operations permitted during repair, and the feedback returned after unsuccessful attempts, allowing the same benchmark to instantiate substantially different repair tasks ranging from localized patch generation to repository-level diagnosis and iterative repair. We present a framework for explicitly specifying the experimental settings associated with reported APR results. We analyze reported experimental settings from systems evaluated on Defects4J and SWE-bench and characterize each result by its task unit, fault-localization assumptions, initial input, tool access, repair-time feedback, final validation, and resource budget. Our analysis shows that benchmark identity alone is insufficient to reconstruct the evaluated task or determine the appropriate scope of comparison across reported repair rates. We therefore introduce a machine-readable schema for specifying each experimental setting to improve reproducibility and make the scope of cross-system comparisons explicit.

Wed 16 SeptSoftware Engineering
The gist
Fixing computer programs automatically with AI models can be done in many ways, but how the repair task is set up affects the results a lot. The authors looked closely at how different studies provide input data, allow tool use, show feedback, and check fixes, finding these factors are often unclear. They created a system that clearly describes these settings so others can understand and compare results better. This helps make research on AI program fixing more repeatable and fair.
Open → 2609.17993v1

Software development adapts itself as projects and conditions change

Rethinking Software Development as a Self-Adaptive Socio-Technical System

Abstract: Agentic AI is expanding across various software engineering activities, yet human-agent organization is typically treated as fixed, which is problematic because development configurations may become suboptimal as development context evolves. We propose viewing software development as a self-adaptive socio-technical system in which both the software project and the development configuration (participants, responsibilities, authority, information, and verification) adapt as goals, evidence, uncertainty, and risk evolve. This yields two coupled forms of adaptation: evolving the software and reconfiguring how subsequent engineering is performed. We illustrate the idea with a proof of concept, outlining challenges and future plans.

Fri 11 SeptSoftware Engineering
The gist
Software development teams usually follow a set plan for who does what, but this plan can stop working well when things change. The authors suggest thinking about software development as a living system that changes both the software and how the team works as new goals, information, and risks appear. This means the people and their roles in the development can adjust alongside the software itself. They provide an initial example of how this idea could work and discuss challenges ahead.
Open → 2609.13461v1

Reinforcement learning method improves code generation tasks at test time

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

Abstract: Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the problem statement, executes candidate programs on these probes, and defines a Probe Consensus Reward (PCR) from the resulting behavioral agreement. PCR provides a behavioral training signal for open-vocabulary programs, but it is not a fully reliable verifier and remains susceptible to reward hacking through spurious consensus. We therefore introduce Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which converts low PCR into conservative negative updates through rank masking and controls policy drift with an entropy ceiling. On coding benchmarks, ERPO substantially improves pass@1 and pass@k in both in-domain adaptation and zero-shot transfer.

Tue 8 SeptMachine LearningComputation and Language
The gist
Generating computer programs is hard because small changes can make big differences, so usual techniques to check answers don't work well. The authors created a way to test programs by running them on special test inputs made from the problem description and seeing if they behave similarly. They use this behavior as a kind of reward to teach the system better. To avoid tricks and incorrect learning, they use a method that carefully adjusts training updates and keeps the program's guessing patterns diverse. Their approach improves the quality of generated code in several tests.
Open → 2609.09135v1