Optimizer choice changes training speed but not data scaling rate
Optimizer-dependent training dynamics converge to the same one-third optimal data scaling
Machine Learning
Summary
Training large language models involves reducing errors as they see more data. The authors show that different training methods (optimizers) affect how quickly the model improves during a single training run but do not change the fundamental relationship between data size and best achievable error rate. They find a consistent scaling pattern where the error drops in proportion to the dataset size raised to the power of one-third, regardless of the optimizer used. This reveals a shared underlying rule in how models learn from data, despite differences in optimization strategies.
What this means in practice
- •For machine learning engineers: Optimize training schedules by selecting and tuning optimizers knowing that data scaling benefits remain consistent across choices.
- •For cloud ai platform operators: Improve resource allocation during large model training by understanding how optimizers affect speed but not sample efficiency.
Authors
Hyunseok Lee, Mihir Basil, Yizhou Liu, Jeff Gore
Abstract
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size $D$. We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to $1/3$ across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents $α_{r}$ and $α_{t}$. Under SGD, both are close to $1/3$, so the data exponent is also $1/3$ across different learning rates. Under Adam the two separate: $α_{r} \simeq 0.48$ but $α_{t} \simeq 0.08$. Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: $D$-independent for SGD but falls with $D$ for Adam. Yet tuned to that optimum, the loss returns to $D^{-1/3}$ for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, $2α_{r}+ α_{t} = 1$, which fixes the optimal data exponent at $1/3$. Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with $D^{-1/3}$ across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.