Policy optimization method avoids unstable action gradients in reinforcement learning
FERPO: Forward Entropy-Regularized Policy Optimization
Machine LearningArtificial IntelligenceRobotics
Summary
When teaching a computer to make decisions in complex situations, one way is to guess how good each action is and improve based on that guess. But predicting the exact value of an action doesn't always help improve decisions correctly because the details can be unreliable. The authors propose a new method called FERPO, which improves decision-making without needing those tricky details by comparing current choices to a carefully balanced target. This approach encourages exploring multiple good options and is more stable and efficient in tests on simulated robotic control tasks.
What this means in practice
- •For robotics software engineers: Train robots to learn movement and control more efficiently by using a policy update method that avoids unstable gradient estimates.
- •For game ai developers: Improve AI agents' decision-making stability and exploration in continuous action spaces without relying on precise value derivatives.
Authors
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Abstract
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).