Reinforcement learning improves by filtering bad training paths

Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning

Machine Learning

Summary

Training AI to make decisions often involves learning from examples called trajectories, but some of these examples can actually hurt learning. The authors show that existing methods to value data don't work well for reinforcement learning because the data is generated on the fly and there's no straightforward way to tell which examples are good or bad. They introduce a new method, Dynamic Trajectory Valuation (DTV), which uses the training process itself to figure out which training paths to keep or ignore, helping the AI learn better and more efficiently. Tests with popular reinforcement learning algorithms show that this approach makes training more stable and improves performance.

What this means in practice

  • For machine learning engineers: Improve stability and data efficiency during reinforcement learning training by filtering out unhelpful trajectories using gradient-based evaluation.
  • For autonomous robotics teams: Enhance online learning of control policies by identifying and ignoring negative experiences that degrade robot behavior during training.

Authors

Xuesong Jia, Ziao Yang, Zhanhe Huang, Hongfu Liu

Abstract

We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.