New method improves language model learning from verbal feedback

Inductive Feedback for Mixed-Policy Distillation

Machine Learning

Summary

Sometimes language models can be taught better by using verbal feedback that points out mistakes and suggests fixes. The authors found that current ways of using this feedback can accidentally transfer the teacher model's own biases or ignore much of the feedback. They developed a new approach that treats feedback as clues for what the next word should be and carefully updates the student model’s learning. Their method also makes better use of the feedback even when the student model doesn’t immediately apply it. Tests show this approach works better on tasks involving knowledge and agent actions.

What this means in practice

  • For ai model developers: Improve training of language models using detailed verbal feedback without requiring automatic verifiers.
  • For conversational ai teams: Enhance dialogue agents by incorporating nuanced corrections and guidance from expert feedback more effectively.

Authors

Amir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang

Abstract

Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.