Momentum helps stabilize high learning rates for faster training progress
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Machine Learning
Summary
Training big language models is tricky because their error landscape looks like a river valley with a low path surrounded by steep sides. The authors study how momentum, a common technique in training, helps keep training stable even with high learning rates that would otherwise cause problems. This stability lets the training move faster along the good low-error paths, improving long-term progress. Interestingly, if the valley is very flat and slow, momentum itself doesn't speed things up directly, but it allows using a bigger learning rate which does.
What this means in practice
- •For machine learning engineers: Optimize training schedules by using momentum to maintain larger learning rates for faster convergence of large models.
- •For deep learning framework developers: Design training tools that implement momentum with stable high learning rates for improved optimization dynamics in language model training.
A theory result. No direct application yet.
Authors
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li
Abstract
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Within such a landscape, gradient descent with large learning rates can move faster along the river despite high apparent loss due to vertical oscillations, while a subsequent sharp decay in the learning rate suppresses these oscillations, revealing genuine optimization progress. This explains the recent success of warmup-stable-decay (WSD) learning rate scheduler which, unlike cosine scheduling, keeps stable high learning rate and decays before producing intermediate checkpoints. Building on this foundation, in this work we take a step further and study the role of momentum within such a loss landscape. We establish theoretical analysis that characterizes how momentum accelerates optimization by stabilizing large learning rates that can not be tolerated by vanilla GD without deviating significantly from the river. The enabled large learning rate in-turn gives greater speed along the river and makes faster essential progress in the long run. Another intriguing observation from theory is that for a river-valley landscape with very flat and slow-spinning river, the momentum itself does not contribute directly to acceleration in terms of the speed of tracking the river, while the main acceleration comes from the admissible larger learning rate.