Adaptive spectral estimation improves matrix optimizers in ai training

Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers

Machine Learning

Summary

Training large AI models often involves adjusting many matrices to speed up learning. Current methods use a fixed routine to process these matrices, which may not be efficient for all parts of the model or training stages. The authors found that by looking at information already computed during these adjustments, it is possible to adapt how the matrices are handled based on their unique properties. This adaptive approach improves accuracy or reduces the work needed, leading to better training results on large models like GPT.

What this means in practice

  • For machine learning engineers: Improve training efficiency and accuracy in AI models by adapting matrix processing to current momentum matrix properties without extra computation.
  • For deep learning framework developers: Integrate spectrum-adaptive orthogonalization routines into matrix optimizers to reduce computation or improve optimizer accuracy in large-scale model training.

Authors

Kristi Topollai, Anna Choromanska

Abstract

Matrix optimizers such as Muon transform each momentum matrix through an approximate orthogonalization, typically implemented by a small number of Newton-Schulz matrix multiplications. The quality and cost of this approximation depend strongly on the singular-value spectrum of its input, yet existing implementations use the same fixed polynomial routine for every layer and throughout training. We show that this uniform treatment is unnecessary: the computations in the Newton-Schulz method already reveal enough information to make the method adaptive. The Gram matrices formed inside Newton-Schulz iterations yield spectral moments through inexpensive scalar reductions, requiring no additional matrix multiplications. From these moments, we recover an estimate of the empirical singular-value distribution and use it to select a polynomial routine specialized to the current matrix. This turns Newton--Schulz orthogonalization into a spectrum-adaptive procedure that responds to differences across both layers and training time. On saved momentum matrices, spectral estimation substantially reduces orthogonalization error at a fixed iteration budget or reaches the same accuracy with fewer iterations, and in GPT pretraining up to 1B parameters it lowers the validation loss of two matrix optimizers. Our results suggest that matrix-function operations inside optimizers need not be designed for a conservative worst-case spectrum: they can cheaply measure the spectrum they are already processing and specialize computations accordingly.