How to Train a Critic Stably and Efficiently
2026-08-24 • Machine Learning
Machine LearningArtificial IntelligenceComputation and Language
AI summaryⓘ
The authors look at a way to improve how AI models learn by using a 'critic' that judges responses during training. They find that previous methods using critics were unstable, so they create a new recipe called Best-Practice Critic Optimization (BPCO) that combines several techniques to make training more reliable. Their method works well on math reasoning tasks across different sized models and can use extra information like grading rubrics during training. The authors show that BPCO matches or beats other methods while needing fewer response samples per prompt.
reinforcement learningcriticadvantage estimationDPPOgeneralized advantage estimationMonte Carlo value targetslanguage modelspolicy optimizationrubric-based rewardmixture of experts
Authors
Penghui Qi, Xiangxin Zhou, Wee Sun Lee
Abstract
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes are often unstable. We study this instability and develop \textbf{Best-Practice Critic Optimization (BPCO)}, a recipe that combines DPPO, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation. Because the critic is used only during training, BPCO can also condition it on reward-defining information, such as a reference answer or grading rubric, that is hidden from the policy. Controlled experiments isolate the effect of each design choice. Across mathematical reasoning tasks with models ranging from 1.5B parameters to 30B-A3B mixtures of experts, BPCO improves a strong critic-based baseline consistently, and matches or exceeds a group-based baseline while sampling one response per prompt. The same recipe also improves learning with rubric-based rewards. These results show that a carefully designed critic provides a reliable alternative to group-relative advantage estimation. Code is available at https://github.com/QPHutu/golden_critic