Papers for

ai development operations teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Explicit preference modeling improves offline alignment of large language models

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

Abstract: Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterative procedure, is a key factor behind its superiority over offline methods. This insight leads us to answer the second question affirmatively and propose Distilled Preference Probability Policy Optimization (DP3O), an effective and efficient offline alignment algorithm. DP3O first learns an explicit preference model using a helper class of LLMs and then distills its knowledge into policy optimization. Theoretically, we show that explicit preference modeling admits better estimation error control than implicit formulations, and that DP3O achieves a tighter generalization bound than hard-label DPO through variance reduction. Empirically, we evaluate DP3O on a wide range of chat-based and downstream tasks and show that it outperforms state-of-the-art offline methods, achieves performance comparable to iterative DPO, and reduces training time by about $42\%$, demonstrating both its effectiveness and efficiency.

Mon 7 SeptComputation and LanguageMachine Learning
The gist
Aligning large language models with human preferences is important but challenging. The authors study why methods that update models in steps usually work better than one-time offline methods. They find that having a clear model of preferences makes a big difference. Using this insight, they create a new method that learns this preference model first, then uses it to train the language model more efficiently. This new approach works as well as stepwise methods but takes less time to train.
Open → 2609.06893v1