Soft actor-critic convergence improved with mirror-descent policy updates
An analysis of Mirror-Descent Soft Actor-Critic
Machine Learning
Summary
Continuous control tasks use a method called Soft Actor-Critic (SAC) to learn actions by balancing reward and exploration. This paper shows how updating the policy using a technique called mirror descent helps prove that the learning will converge to a good solution. The authors find mathematical conditions that ensure stable training depending on how the value estimates curve. They also show how this method controls errors better than previous ways, which helps in more reliable learning for complex tasks.
What this means in practice
- •For robotics engineers: Improve convergence guarantees when training robots using continuous control reinforcement learning algorithms.
- •For autonomous vehicle developers: Develop more stable reinforcement learning models for vehicle control by leveraging mirror-descent policy updates.
A theory result. No direct application yet.
Authors
Denis Zorba, Michal Valko
Abstract
Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor objective, characterised by the curvature of the $Q$-function estimate through the Legendre differential operator, and establish an $\mathcal{O}\!\left(N^{-\frac{1}{5}}\right)$ best-iterate finite-time convergence rate up to actor and critic approximation errors. Moreover, the mirror-descent step size $λ$ directly controls the target drift and hence actor tracking error, whereas the analogous Gibbs bound contains a non-vanishing tracking term.