Bias adjustment improves fairness in language model preference training

BA-DPO: Bias-Adjusted Direct Preference Optimization for Language Model Alignment

Machine LearningArtificial Intelligence

Summary

When training language models to prefer better answers, human judges can bring in their own hidden biases, like favoring longer responses or certain styles. The authors introduce a new method called Bias-Adjusted DPO that helps the model learn to ignore these biases by figuring out each judge's preferences for specific traits. This adjustment makes the model’s training fairer and produces better balanced results without losing quality. Their tests show it can significantly reduce bias effects on model outputs.

What this means in practice

  • For ai model trainers: Train language models to produce responses less influenced by human annotator biases across various attributes.
  • For ai ethics teams: Adjust model outputs to reduce unintended bias amplification during alignment for fairer AI interactions.

Authors

Antonio Ferrara, Alberto Rumi, Francesco Bonchi

Abstract

Preference-based alignment methods such as Direct Preference Optimization (DPO) use pairwise preferences labeled by human annotators to fine-tune language models. However, annotators carry systematic biases toward some attributes: a name that signals a gender or an ethnicity, a persona, a language variety, a formatting convention, or length. If not properly addressed, these systematic biases can be absorbed and amplified during alignment. Existing methods address length bias or annotator disagreement, but fail to eliminate biases toward arbitrary attributes. To address this limitation, we propose Bias-Adjusted DPO (BA-DPO), a generalization of DPO that adds one bias parameter per annotator toward responses carrying a declared attribute. We prove that the objective is convex in the bias parameters and that the votes identify each annotator's bias up to a shared constant. The remaining constant is what fixes the aligned model's attribute rate: by default the rate of the reference model, or a target rate, which we use to bring a biased policy to statistical parity. On a corpus with planted biases, DPO drives the attribute from a balanced start to probability 0.96 and BA-DPO removes 81 to 95\% of that shift; on MultiPref with real annotators it removes about half of DPO's lengthening. Both hold at 0.5B with full fine-tuning and at 8B with LoRA, at no higher KL than DPO and no loss in judged quality.