Papers for

software reliability engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Benchmark evaluates AI agents configuring software deployment environments

FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

Abstract: Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes. Agents submit declarative artifacts that are collected, rebuilt, and redeployed in a pristine environment. Four gated binary check layers measure build, readiness, behavior, and conformance to the deployment specification, using programmatic checks without an LLM judge. A four-arm release gate requires a resolving reference solution and rejects tasks solved by do-nothing, specification-transcription, or generic-stub submissions. The released check annotations expose the link between 2,145 checks and their specifications, including seven documented gaps. Three additional adversarial strategies test shortcuts in the grading signals; none resolves any of the 135 tasks they cover, while a vacuous health probe passes readiness and exposes the need for downstream checks. On the 136-task evaluation grid, seven language models from four providers use the same four-tool scaffold and resolve 52.9-75.0 percent of tasks. The three zero-intelligence floors resolve none and reach a mean Deployment Score of at most 0.44. Readiness is the largest failure stage, accounting for 110 of 313 unresolved episodes. Mean resolution rate is 30.7 percentage points higher on the repair task group than on the disjoint greenfield group, with a positive gap for every model; ten tasks resist all seven. In a 25-task case study, one practicing engineer directing Claude-Sonnet-5 resolves 92 percent against 72 percent for the autonomous baseline. FDE-Bench links deployment success and failure to artifacts that can be inspected and replayed.

Wed 23 SeptSoftware EngineeringArtificial Intelligence
The gist
Setting up software so it runs properly and stays working can be tricky. The authors created a test suite called FDE-Bench that checks how well AI systems can turn code into working software setups using technologies like Docker and Kubernetes. Their tests include building, starting, and checking if the software works as expected, without using AI to judge the results. They found AI can solve many but not all deployment tasks, especially struggling with readiness checks. The benchmark also helps engineers understand and improve deployment processes by linking successes and failures to specific configuration artifacts.
Open → 2609.27571v1

System fixes broken traces across video platform microservices

Backstitch: Restoring Request Causality Across a Production Microservice Fleet

Abstract: A major video platform runs on thousands of microservices, each request propagating a context so downstream work can be traced and governed. At handoffs outside instrumented paths, e.g., custom queues and callbacks, the payload continues but the context does not, and the request still succeeds under existing tests. Such breaks are silent and widespread: 673 of 1,133 services carried at least one. Backstitch, a specialized agentic system, repairs them using the surviving execution as its reference: replay determines whether a suspicious call is request-correlated, source analysis reaches the responsible handoff, a bounded change restores its contract, and the same replay validates the fix. Repairs restore the causal chain without disturbing the work it describes: breaks at 240 of the repaired calls fell from 90.46% to 4.69%, and over 112 days the fleet's break rate more than halved.

Wed 23 SeptDistributed, Parallel, and Cluster Computing
The gist
Many small services handle video requests, passing information along so the whole process is understood. Sometimes the connection breaks silently, causing loss of tracking without failing the request. The authors created Backstitch, a system that finds where these breaks happen and repairs them automatically by analyzing and replaying the service calls. This restores the tracking without changing the actual work, cutting errors in half over months.
Open → 2609.27538v1

Graph structure adds little to microservice error detection accuracy

Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring

Abstract: Graph neural networks dominate recent work on microservice root-cause analysis, yet recent results question whether the graph contributes. Those results compare whole pipelines, so when a flat model wins one cannot tell whether structure is useless or redundant. We run the comparison they imply on RCAEval: three learned arms with identical features, optimiser, validation split, early-stopping rule and scoring head, in which a single term separates the graph arms. Across two RCAEval benchmarks, two topology sources and four regimes we find no reliable graph-specific effect: in-distribution the graph model leads the conventional flat model by 0.003 Avg@5 (p = 0.844, n = 6 disjoint folds). Auditing the pipeline surfaced two benchmark properties that condition any result on it. RCAEval injects faults into only five services per system while exposing 12 to 70 in telemetry, and the headline metric is Avg@5: a ranker reading no telemetry at all places the true culprit in the top five on 99.7 percent of held-out incidents, scoring Avg@5 0.488. That prior, not the uniform-random 0.137, is the honest in-distribution floor, and it collapses to 0.192 across systems. The second property is a non-uniform column schema that silently zeroes telemetry for most RE1 cases. We reproduce a published baseline, BARO, RCAEval's own reference implementation; on the one system with a clean schema it reaches similar aggregate accuracy to our heuristic, within 0.004, under a different scoring rule. The audit motivated a new model. PSC-GRCA separates a candidate score into a system prior, telemetry evidence and a centred graph residual, and reaches mean Avg@5 0.915 against 0.864 for the flat baseline, while its ablations locate most of the gain in the prior term rather than the graph. We close with a twelve-item checklist for graph-versus-flat ablation studies, distilled from sixty-two recorded defects.

Tue 22 SeptSoftware EngineeringMachine Learning
The gist
Finding the root cause of problems in complex software made of many small parts is hard. The paper studied whether using graph-based models, which consider how parts connect, improves this process compared to flat models that ignore connections. They found that graph models only slightly outperform flat models, and much of the advantage comes from prior knowledge rather than the graph itself. The study also highlights issues in common testing setups that might overstate graph benefits. The authors suggest ways to better evaluate graph use in such tasks.
Open → 2609.27069v1

Agentic AI generates reliable Linux utility programs with human oversight

A Study of the Reliability of Agentic AI-Generated Programs

Abstract: Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question based on three practices. First, we applied a typical best-practices agentic AI workflow for software development. Second, our target programs were ten well-known, release-quality human-written Linux utility programs so that we could compare the AI-generated code against a concrete ground truth. Third, we based our measure of reliability on a widely used testing technique, fuzz random testing. For this testing, we used both classic black box, generational testing and more modern coverage guided (gray box, mutational) testing using AFL++. We found that the AI-generated versions of the utility programs were typically as reliable - often more reliable - than the latest human-generated versions of these programs. While the AI-generated versions did have some failures, they were less common than the code from the standard repositories. Interestingly, the AI-generated code was less likely to have failures such as memory errors (such as buffer overflows) but more likely to have hangs such as infinite loops. In addition, we verified that generating robust and reliable software using agentic AI requires careful practice and human supervision. The quality of the code is highly dependent on the prompts and skills used, and how the human directing the process responds. We also demonstrated that using agentic AI workflow for software development (with its prompts and skills) can become a specification of the code that leads to cost-effective sustainability of the software.

Wed 16 SeptSoftware EngineeringArtificial Intelligence
The gist
This study looked at whether programs written by AI can be trusted as much as those written by people. The researchers asked AI tools to create ten well-known Linux utility programs and then tested these AI-made programs with special software tests. They found that the AI-generated code was usually just as reliable or even better than human code, having fewer memory problems but being more prone to getting stuck in loops. However, the quality of AI code depends a lot on how skilled the human guides the AI and what instructions they give. The study shows AI can help make solid software if humans carefully watch and guide the process.
Open → 2609.18298v1