Finite-time convergence rates for two-speed stochastic learning with constraints
Finite-Time Concentration and Convergence Rates for Projected Two-Time-Scale Stochastic Approximation with Markov Noise
Machine Learning
Summary
This paper looks at a mathematical method for learning systems that update at two different speeds under some limits and randomness influenced by changing conditions. The authors show how to measure how fast these updates get close to their goals even when the system has boundaries and unpredictable noise. They provide clear ways to separate different sources of error and prove the method will reliably settle down over time, improving certain learning algorithms like actor-critic methods. Their results also apply to other constrained learning tasks that use projection and random feedback.
What this means in practice
- •For reinforcement learning engineers: Improve reliability and convergence speed in actor-critic algorithms with constrained policy classes using the paper’s finite-time error bounds.
- •For machine learning practitioners: Apply the convergence guarantees to projected TD(0) and stochastic gradient descent methods under Markov noise and boundary constraints.
Authors
Rahul Singh, Vivek S. Borkar, Eric Moulines
Abstract
We study finite-time concentration and convergence rates for projected two-time-scale stochastic approximation driven by a controlled Markov chain. The averaged fast map is contractive, while the slow iterate is projected onto a compact convex polyhedron. The associated projected ordinary differential equation may have a discontinuous vector field at the boundary, preventing a direct application of standard analyses based on Lipschitz vector fields. Using the Skorokhod map, we establish explicit high-probability bounds for tracking the moving fast equilibrium and the projected slow dynamics. These bounds separate martingale fluctuations, Markov-noise residuals, and the bias due to time-scale separation. A Lipschitz Lyapunov function satisfying a uniform decrease condition over fixed time intervals yields almost-sure convergence, with explicit last-iterate rates when the decrease admits a power lower bound. Under uniform Lyapunov contraction, polynomial step sizes yield joint fast-tracking and slow Lyapunov-error exponents arbitrarily close to $1/3$. Under the additional assumption that the reduced slow update map is a Euclidean contraction, logarithmically separated step sizes improve the joint rate to $O(n^{-1/2}\log n)$ almost surely, including for boundary equilibria. The same rate holds under a distinct geometric condition involving a strictly attracting face of a box and a fast equilibrium that is constant on that face. An actor-critic application achieves an almost-sure value-gap rate of $O(n^{-1}\log n)$ relative to the optimum within the constrained policy class. Further applications include projected TD(0) and projected stochastic gradient descent. We also extend the analysis to projection of the fast recursion under Euclidean contractivity.