Efficient self-distillation improves safety and reasoning in large language models
EOPSA: Efficient On-Policy Self-Distilled Safety Alignment
Artificial Intelligence
Summary
Training language models to be safer and better at reasoning can be inefficient because the teacher model guiding the student becomes less helpful the longer the student generates text. Also, teaching signals get confused by style differences that don't relate to safety. The authors identified these problems and created a new method called EOPSA that focuses the training on key parts of the text where safety matters most and only when the teacher’s guidance is reliable. This approach reduces computation by half and improves both safety and reasoning in large models.
What this means in practice
- •For ai safety engineers: Train language models with improved safety measures using more efficient self-distillation techniques focused on critical safety tokens.
- •For natural language processing teams: Enhance large-scale language model training efficiency and reasoning retention by selectively backpropagating gradients on safety-relevant outputs.
- •For chatbot developers: Build safer conversational agents by integrating targeted distillation methods that reduce computation while improving output compliance with safety criteria.$Commercial implications: Enables creation of safer commercial chatbots with efficient training pipelines enhancing compliance and reducing costs.
Authors
Qirui Liu, Yichen Sun, Yan Wang, Yu Mi, Wei Cao, Yue Shen, Zhixuan Chu, Kui Ren
Abstract
On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher's corrective efficacy degrades precipitously as the student's generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by $\sim$50% and backpropagates through merely $\sim$2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.