MeqMuon optimizer improves training efficiency for large language models
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
Machine Learning
Summary
Training very large language models requires a lot of computing power and memory. The authors propose MeqMuon, an improved method for tuning these models during training, which balances different parts of the training updates more effectively. This method also uses less memory by avoiding storage of some extra calculations needed in previous methods. Experiments show MeqMuon leads to faster and better training compared to existing approaches.
What this means in practice
- •For machine learning engineers: Reduce training time and memory usage when pretraining large language models by integrating MeqMuon optimizer.
- •For cloud infrastructure teams: Optimize resource allocation and lower costs for large-scale language model pretraining using MeqMuon’s efficient memory usage.
Authors
Chang-Wei Shi, Xu Wang, Wu-Jun Li
Abstract
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.