Papers for

software test engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Large language models improved by fixing math reasoning steps

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Abstract: While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.

Thu 1 OctMachine Learning
The gist
Solving math problems with AI is tricky because it needs clear step-by-step understanding, not just guesses. The authors studied what kinds of math reasoning large language models (LLMs) can and cannot do well, finding that discovering new math ideas is the hardest part. They created a test to measure different reasoning skills and found ways to fix common mistakes after training. Their new method helps models learn better math reasoning by focusing on these key reasoning parts, improving their problem-solving skills.
Open → 2610.02191v1

Agent error dataset enables better failure diagnosis and correction

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

Abstract: An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.

Wed 30 SeptArtificial IntelligenceComputation and Language
The gist
Figuring out where AI agents make mistakes is tricky but valuable. The authors created a huge dataset of over 50,000 cases where AI agents erred, paired with explanations and suggested fixes. They showed that using this dataset to train AI models improves their ability to diagnose errors and propose better actions, making the AI agents more reliable. This work helps AI systems learn from their failures more effectively.
Open → 2609.40111v1

Benchmark reveals how communication shapes multi-agent system performance

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

Abstract: Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.

Fri 18 SeptMachine Learning
The gist
Coordinating multiple AI agents involves graphs that dictate how they share information and split tasks, but it's hard to tell which part of this setup affects their success. The authors created OpenMAS-GCom, a test environment that changes one piece of the system at a time to pinpoint how communication links, roles, and information flows impact performance. They tested many configurations on diverse tasks, discovering that losing specialist roles hurts more than losing critics, and that wrong messages or worker failures degrade performance differently. This benchmark helps understand what really matters in multi-agent communication setups.
Open → 2609.21527v1