A framework improves reinforcement learning efficiency using better prompt evaluation

MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR

Machine Learning

Summary

Reinforcement learning can help language models get better by learning from feedback, but it can be slow and costly. The authors found that existing methods for choosing which prompts to use during training suffer from unpredictable noise that limits accuracy. They developed a new approach called MaPP that reduces this noise by mathematically accounting for uncertainty in responses. This method leads to more reliable learning signals and smarter prompt choices, helping models improve faster on tasks like math and planning.

What this means in practice

  • For ai engineers: Improve training efficiency of language models by using MaPP for better prompt prioritization during reinforcement learning.
  • For machine learning platform teams: Integrate MaPP’s posterior predictive method to reduce rollout costs and improve accuracy in reinforcement learning model tuning.

Authors

Yangyang Ren, Haodong Zhu, Sheng Xu, Yanjing Li, Nikolai Yu. Zolotykh, Wentao Zhang, Baochang Zhang

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.