Matrix adaptive optimization improves neural network training stability
Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
Machine LearningArtificial Intelligence
Summary
Optimizing neural networks often involves adjusting many parameters arranged as matrices. Typical methods adapt learning rates for each parameter individually, missing patterns in the matrix structure. The authors develop a new approach that adapts learning rates by looking at entire rows or columns of these matrices, making training more stable and effective. Their method shows better performance when parameters have structured gradients and helps train deeper or larger neural networks.
What this means in practice
- •For deep learning engineers: Enable more stable training of deep neural networks by adapting learning rates along matrix rows or columns to better match parameter structures.
- •For recommendation system developers: Improve matrix factorization techniques by using row-wise or column-wise adaptive scaling to accelerate convergence and increase optimization stability.
Authors
Wenpeng Zhang, Runsheng Yu, Peilin Zhao
Abstract
Adaptive optimization methods such as AdaGrad and Adam are widely used in modern neural-network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop a general Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled approach to deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions and analyzing the resulting regret trade-off, we derive Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), with adaptive scaling determined by the accumulated row-wise or column-wise gradient norms. We establish regret guarantees and show that these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training further demonstrate the benefits of aligning adaptive scaling with matrix structure, including improved optimization stability and trainability at larger learning rates and greater network depths.