Fisher-informed method improves stability of feedback learning in large language models
Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
Machine Learning
Summary
Training large language models to learn from their own outputs can be unstable and cause performance to collapse. The authors propose a method called FIRE that helps the model better handle feedback by separating how feedback moves the model from how far it moves. This method recalibrates training signals based on whether answers are correct or incorrect, using information from the model’s internal statistics. Their experiments show that FIRE makes training more stable while keeping the model’s performance strong.
What this means in practice
- •For ai model trainers: Improve stability when fine-tuning large language models using feedback from their own outputs to prevent performance collapse.
- •For machine learning engineers: Use recalibrated feedback supervision to enhance model updates and maintain strong performance during on-policy self-distillation.
- •For automated customer service teams: Enhance chatbot learning from interaction feedback to produce more reliable responses over time through stable fine-tuning.$Commercial implications: Enables development of more consistent chatbots by improving feedback-based learning stability, increasing product quality for customer support solutions.
Authors
Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. Brinton
Abstract
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.