Papers for

software automation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Self-explaining language models improve task solving without reinforcement learning

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Abstract: People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.

Mon 28 SeptArtificial IntelligenceComputation and Language
The gist
People learn not only by doing but also by thinking about and explaining their experiences, which helps them improve. The paper shows that a language model can train itself to do better on tasks just by explaining what it did, without needing rewards or teachers. This method, called ROFT, made the model solve more problems and even learn from tasks where it initially failed every attempt. The explanations help the model figure out which actions were good or bad, leading to better future behavior. This suggests teaching AI to explain could help it learn more effectively.
Open → 2609.35741v1

Hexis organizes agent skills into clear step-by-step machines

HEXIS: Compiling Skills into Extended Finite State Machines

Abstract: Agent skills provide reusable knowledge and instructions, yet agents must repeatedly infer how to apply them and which operation should follow. This couples task reasoning with control decisions, allowing prescribed steps to be omitted or applied incorrectly. We introduce HEXIS, which compiles agent skills into extended finite state machines that separate knowledge from control flow. Skill knowledge is incorporated into local instructions that guide reasoning and generation within states. The machine records execution progress and intermediate results, while explicit transition conditions determine subsequent operations. Our incremental compiler first maps skill clauses and tool interfaces to state operations, local instructions, data bindings, and transitions. It then aligns development traces with existing states to identify missing operations and dependencies. These are incorporated by adding or reusing states and refining their connections. Updates are accepted only after static checks and replay of the current and all previously accepted traces. Across four benchmarks and four executors, HEXIS improves success over Skill + ReAct by 16.1 percentage points on average. Qwen3.8-27B reduces execution tokens by 38.4-88.9% across benchmarks.

Thu 24 SeptArtificial Intelligence
The gist
Making software agents follow complex tasks is hard because they often mix deciding what to do with how to do it, causing mistakes. The authors propose HEXIS, which turns skills into a machine that separates knowledge from control steps, helping agents know exactly what to do next. This makes tasks run more correctly and efficiently, as shown in tests where HEXIS performed better than older methods and used fewer computing steps.
Open → 2609.30123v1

Policy editing improves workflow agents by limiting harmful ripple effects

Local Edits, Global Ripples: Replay-Informed Policy Adaptation for Workflow Synthesis

Abstract: Prompt-policy editing offers a practical way to improve agents that synthesize executable workflows without updating the underlying model. However, persistent prompt editing has two coupled properties. First, edit locality does not imply effect locality: an edit confined to one policy segment can ripple through downstream execution, altering behavior beyond the edited segment. Second, edit effects are composition-sensitive: edits that work in isolation can interfere after composition, causing one or both to lose their benefit or become harmful. Persistent adaptation must therefore support two distinct decisions: identifying where the policy should change from execution feedback, and determining whether the resulting edit remains safe to persist after composition. To address these challenges, we introduce RIPPLE (Replay-Informed Persistent Policy Localization and Editing), which separates where an edit is made from whether it remains safe after composition. It diagnoses failed trajectories, maps each actionable failure to a predefined policy segment, and restricts the correction to that part of the policy. RIPPLE then evaluates candidates against the same iteration-start policy to compare their isolated gains, before replaying promising edits after previously accepted updates to expose downstream effects and interactions. Only edits that remain safe under composition are retained. We evaluate RIPPLE on Flow-HO, a synthetic held-out benchmark for executable workflow synthesis. RIPPLE improves validation success by up to 23.1% and yields positive gains on two additional frozen language-model backbones, while maintaining edit efficiency and low execution cost. Targeted interaction analysis further demonstrates both properties: a segment-local tool-use edit changes downstream resource resolution and validation, while an edit beneficial in isolation becomes harmful after composition.

Thu 10 SeptComputation and LanguageSoftware Engineering
The gist
When computer programs automatically create step-by-step workflows, changing one part can unexpectedly affect later parts, sometimes causing problems. The paper presents RIPPLE, a method that carefully fixes only the parts that fail and checks if those fixes cause trouble elsewhere before keeping them. This helps the programs improve without causing new errors. The authors tested RIPPLE on a task where a computer builds workflows, and it made the programs work better by up to 23%.
Open → 2609.12127v1