Decoupling credit direction and magnitude improves self-distillation in ai reasoning
Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
Artificial Intelligence
Summary
Teaching AI models to learn from themselves can be tricky because it's hard to tell which steps in their reasoning deserve credit and by how much. The authors found that mixing these two (direction and amount of credit) with teacher guidance can cause errors and confusion. They propose a new way to separate these two signals so the AI can better learn from its own experience and correct mistakes in guidance. Their approach improves performance on math and multimodal reasoning tasks across many tests.
What this means in practice
- •For ai model developers: Improve reasoning accuracy in language and multimodal AI models by applying decoupled credit signals during self-distillation training.
- •For natural language processing engineers: Enhance token-by-token credit assignment for policy optimization to boost performance on complex reasoning benchmarks.
Authors
Yugu Li, Zehong Cao, Peizhen Li, Yang Zhang, Siyi Hu, Jianglin Qiao
Abstract
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5$\times$ reduction in token credit magnitude.