Combining sparse rewards and dense guidance can destabilize learning
When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation
Machine Learning
Summary
When training AI agents with only a simple win-or-lose signal, it can be hard for the agent to learn which actions were good or bad. To help, people add detailed hints from a teacher’s advice to guide the agent’s learning on every step. This paper studies what happens when these two guidance signals don’t align well during training, causing the process to become unstable or fail. The authors analyze the interactions and identify conditions that cause problems, then propose new techniques to balance these signals and improve training stability.
What this means in practice
- •For machine learning engineers: Improve training stability for reinforcement learning agents combining sparse rewards with teacher guidance across diverse environments.
- •For natural language processing developers: Design better on-policy training algorithms for language models that integrate sequence-level rewards and token-level teacher signals.
Authors
Xinke Jiang, Tao Feng, Zhibang Yang, Zhixin Zhang, Weixuan Xu, Haoyu Zhang, Xu Chu
Abstract
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a scalar-weighted teacher KL term to the policy-gradient objective, providing dense token-level guidance that may be unreliable at some positions. Despite the benefits of combining these signals, their interaction during optimization can destabilize joint training. To understand how this instability develops, we study the learning dynamics of hybrid reward--distillation training through a neural tangent kernel (NTK) analysis. We introduce the cross-signal NTK $K_{DR}(n)$, a token-level statistic that measures the alignment between reward and distillation gradients at position n. Through this analysis, we identify two failure modes: 1 Magnitude drowning, where the reward gradient exceeds the distillation gradient by orders of magnitude, so that even weak directional conflict can cause the distillation loss to rise despite its explicit inclusion in the training objective; and 2 Localized directional conflict, where the sequence-level advantage and the teacher's position-specific distribution induce opposing updates at the same token ($K_{DR}(n)\!<\!0$). The severity of these effects depends on the optimization regime: the gradient-norm ratio $κ\!=\!\|\nabla\mathcal{L}_R\|/\|\nabla\mathcal{L}_D\|$ varies by roughly an order of magnitude across tasks, and our experiments reveal an empirical threshold beyond which naive mixing can lead to persistent training collapse. Motivated by these findings, we introduce the M3 family, which combines magnitude normalization with three strategies...