Explicit preference modeling improves offline alignment of large language models
Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation
Computation and LanguageMachine Learning
Summary
Aligning large language models with human preferences is important but challenging. The authors study why methods that update models in steps usually work better than one-time offline methods. They find that having a clear model of preferences makes a big difference. Using this insight, they create a new method that learns this preference model first, then uses it to train the language model more efficiently. This new approach works as well as stepwise methods but takes less time to train.
What this means in practice
- •For ai model training teams: Improve training efficiency and performance of large language models using offline preference distillation approaches.
- •For ai development operations teams: Reduce training time for aligned language models by adopting methods that combine explicit preference models with policy optimization.
Authors
Wenbo Zhang, Wenzhuo Zhou, Hengrui Cai, Zhengling Qi
Abstract
Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propose Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm. DP3O first learns an explicit preference model using a helper class of LLMs and then distills its knowledge into policy optimization. Theoretically, we show that explicit preference modeling admits better estimation error control than implicit formulations, and that DP3O achieves a tighter generalization bound than hard-label DPO through variance reduction. Empirically, we evaluate DP3O on a wide range of chat-based and downstream tasks and show that it outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time by about $42\%$, demonstrating both its effectiveness and efficiency.