Rubric response theory improves reward scoring from partial rubric feedback

Rubric Rewards from Item Response Theory

Computation and Language

Summary

Many tasks for AI, like grading language answers, don’t have a single correct answer. Instead, experts use rubrics with criteria to judge quality, but combining those into a single score can be tricky and inefficient. The authors propose Rubric Response Theory (RRT), a new way to turn rubric judgments into a more accurate and adaptive reward score by modeling each criterion’s difficulty and importance. This method reduces how many judgments are needed while improving scoring compared to older methods.

What this means in practice

  • For ai system trainers: Improve scoring accuracy and reduce human judge requests when training language models on open-ended tasks using rubrics.
  • For language assessment teams: Generate more reliable quality scores from partial rubric data for evaluating language proficiency or student responses.

Authors

Milad Yazdani, Yaser Souri, Xiren Zhou, Pranit Chawla, Dena Shahriari, Subhojit Som, Xia Song

Abstract

Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.