Negative self-distillation improves large language model reasoning skills

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Computation and LanguageMachine Learning

Summary

Using current methods where a language model copies its own confident but possibly flawed answers can actually hurt its reasoning ability. The authors show that having models learn by avoiding their own mistakes, instead of imitating perfect answers, can lead to better reasoning. They created a way to identify only the parts of an answer that show errors and teach the model to avoid those while keeping its general language skills intact. Tests show this new method consistently works better than previous ones for reasoning tasks.

What this means in practice

  • For ai development teams: Train language models to improve complex reasoning by learning to avoid flawed logic rather than mimic confident answers.
  • For automated customer support teams: Enhance chatbot reasoning to reduce confidently wrong responses by teaching models to recognize and avoid error patterns.

Authors

Rongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu, Wei-Lin Chen, Yu Meng

Abstract

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.