Muon changes learning stability patterns in large language model training

Muon Sublates the Edge of Stability in LLM Pretraining

Machine Learning

Summary

Training large language models like those used in AI often involves carefully balancing how fast the model learns and how stable that learning is. The researchers found that when using a training method called Muon, the usual signals that tell us when the training is at a critical balance point don’t behave as expected. They showed that Muon separates the loss balance and the direction alignment in training, meaning these two aspects respond differently to training settings. This insight helps better understand and potentially improve how large AI models learn.

What this means in practice

  • For machine learning engineers: Optimize training schedules for large language models by considering Muon's unique stability dynamics to improve learning efficiency.
  • For ai infrastructure teams: Design resource allocation strategies that account for differing responses of loss and alignment to batch size and learning rate in Muon-based training.

Authors

Yanzhe Chen, Qifang Zhao, Xiaoxiao Xu, Fanghui Liu

Abstract

Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary $2ρ_b/η$, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining