How learning rates transfer in shallow networks with longer training times

Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons

Machine Learning

Summary

Tuning the learning rate is important when training neural networks, but it can be costly for large networks. This paper studies how to transfer good learning rates from smaller to larger shallow linear networks when the training time grows with the size of the network. The authors find conditions under which this transfer works well and explain how the data's spectral properties affect it. They also describe how the extreme data features influence the best learning rate and final training quality.

What this means in practice

  • For machine learning engineers: Tune learning rates on smaller networks and transfer them reliably to larger models during long training cycles to reduce hyperparameter search cost.
  • For data scientists in signal processing: Use the understanding of spectral effects on learning rate transfer to better tune linear models handling large structured datasets.

A theory result. No direct application yet.

Authors

Mana Sakai, Masaaki Imaizumi

Abstract

Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as $n,T\to\infty$ whenever $T=o(\sqrt{n})$. (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.