Papers for

ai systems engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Self-explaining language models improve task solving without reinforcement learning

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Abstract: People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.

Mon 28 SeptArtificial IntelligenceComputation and Language
The gist
People learn not only by doing but also by thinking about and explaining their experiences, which helps them improve. The paper shows that a language model can train itself to do better on tasks just by explaining what it did, without needing rewards or teachers. This method, called ROFT, made the model solve more problems and even learn from tasks where it initially failed every attempt. The explanations help the model figure out which actions were good or bad, leading to better future behavior. This suggests teaching AI to explain could help it learn more effectively.
Open → 2609.35741v1

Large reasoning models safety depends on first generated token

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited understanding of the internal mechanisms behind safety failures. In this work, we investigate this failure through a token-level positional analysis of refusal dynamics and identify a localized vulnerability at the onset of reasoning, which we term Onset Refusal Collapse (ORC). We find that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation. Motivated by this finding, we propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset. Despite updating only a single token embedding, SafeToken effectively mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. These results suggest that safety failures in LRMs can arise from a transient breakdown at the critical transition from understanding to generation.

Wed 16 SeptArtificial Intelligence
The gist
Large reasoning models can solve hard problems but sometimes fail to refuse harmful questions safely. The authors found that these safety failures happen right when the model starts generating an answer. They call this problem Onset Refusal Collapse (ORC). To fix this, they created SafeToken, which adjusts just the first token during response generation to keep the model safer without hurting its problem-solving skills.
Open → 2609.18471v1