Persistent negatives stabilize reward training in black-box policy learning
Persistent Negatives for Adversarial Black-Box On-Policy Distillation
Computation and LanguageArtificial Intelligence
Summary
Training AI systems by learning from their own generated responses can be tricky when only sample answers are available, not detailed probability info. The paper studies a way to improve this by using a technique that compares current AI responses to both recent and past examples, preventing the training from chasing a moving target. This helps the system learn more accurately and consistently. The authors show that using a mix of new and historical comparisons improves performance and stability across various tests.
What this means in practice
- •For chatbot developers: Improve chatbot response quality by stabilizing reward signals during model training without access to internal probabilities.
- •For ai system trainers: Enhance training of AI policies using sample-only teacher feedback by anchoring discriminator rewards with past comparisons.
Authors
Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu, Peggy Yang, Dongzhuo Li, Ruiyi Li, Serena Li, Gedi Zhou, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Lizhu Zhang
Abstract
Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator compute, historical comparisons train the discriminator, while GRPO remains on-policy with fresh student responses. Our analysis identifies the Bayes-optimal reward as a teacher-to-negative log-density ratio and, under explicit assumptions, shows how persistent negatives anchor the discriminator and reduce reward-estimation MSE relative to fresh-negative training. Across two student families, three judges, and four judged-chat benchmarks, persistent-negative adversarial distillation consistently improves performance over current methods at matched discriminator compute. It also yields smoother fresh-policy discriminator trajectories, with fewer below-chance dips. These findings identify the discriminator's negative distribution as an important design axis in black-box on-policy distillation.