Papers for

automated content moderators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Large reasoning models safety depends on first generated token

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited understanding of the internal mechanisms behind safety failures. In this work, we investigate this failure through a token-level positional analysis of refusal dynamics and identify a localized vulnerability at the onset of reasoning, which we term Onset Refusal Collapse (ORC). We find that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation. Motivated by this finding, we propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset. Despite updating only a single token embedding, SafeToken effectively mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. These results suggest that safety failures in LRMs can arise from a transient breakdown at the critical transition from understanding to generation.

Wed 16 SeptArtificial Intelligence
The gist
Large reasoning models can solve hard problems but sometimes fail to refuse harmful questions safely. The authors found that these safety failures happen right when the model starts generating an answer. They call this problem Onset Refusal Collapse (ORC). To fix this, they created SafeToken, which adjusts just the first token during response generation to keep the model safer without hurting its problem-solving skills.
Open → 2609.18471v1