Adaptive bias control improves language model alignment with limited human data
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
Machine LearningArtificial Intelligence
Summary
Training language models to follow human preferences usually requires a lot of expensive human feedback. The authors found that using AI-generated labels is easier but introduces errors. Their method, ABC-Align, smartly corrects these errors by using a small amount of human-labeled data to guide adjustments, making the training more accurate. They showed this approach works better than previous methods when human feedback is limited.
What this means in practice
- •For language model developers: Enhance training of large language models to better match human preferences when human feedback is scarce.
- •For machine learning engineers: Reduce training instability caused by biased pseudo labels using adaptive bias correction in semi-supervised setups.
Authors
Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff
Abstract
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at https://github.com/SewoongLab/abc-align .