Policy optimization method improves learning from reused reinforcement data
ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
Machine LearningArtificial Intelligence
Summary
Reinforcement learning methods often reuse past experiences to improve a policy, but this can cause problems when the new policy behaves differently from the data it learned from. The authors describe a specific problem where some important positive learning signals get ignored because of how gradients are clipped during training. They propose a new method called ReSPO that adjusts how the learning updates are shaped, allowing the model to better learn from rare positive examples and reduce the impact of misleading negative examples. This helps the model learn faster and perform better on tests using reused data.
What this means in practice
- •For machine learning engineers: Improve reinforcement learning algorithms that reuse training data to better learn from rare but important experiences and reduce harmful updates.
- •For ai model trainers: Enhance training of large models with reinforcement learning by enabling more effective use of long positive reward sequences during early optimization stages.
Authors
Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh
Abstract
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated positive responses at the low-importance-weight tail while permitting severely over-generated negative responses to dominate the high-weight tail. To address this, we propose ReSPO (Reshaped Sequence Policy Optimization), which replaces clipping with a smooth, two-branch sequence-level kernel derived from an $α$-divergence variational objective and an exponential variance-control tilt. The positive branch preserves a nonzero gradient weight for under-generated positive responses, while the negative branch suppresses heavily over-generated negative responses. We demonstrate that ReSPO effectively learns from long positive reasoning trajectories during early training, even when accumulated policy drift relegates them to the low-importance-weight tail. On dense and MoE Qwen3 models, ReSPO accelerates early optimization, improves final training scores, and achieves higher held-out benchmark performance under a rollout reuse, validating our approach on importance-weight tail control in off-policy learning.