Adam optimizer behavior revealed through a stable ratio structure

The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics

Machine LearningArtificial Intelligence

Summary

Adam is a popular tool used to help train AI models, but how it adjusts itself over time is complex and not well understood. The authors studied a version where two main settings are equal and found that Adam’s behavior can be explained by looking at a special ratio that stays stable across different tasks and models. This insight helps simplify how Adam stores information, allowing it to be compressed without losing performance. They also showed how this view links Adam to simpler methods that only use the sign of updates, providing clearer ways to switch between them.

What this means in practice

  • For machine learning engineers: Store optimizer state efficiently using low-bit compression while keeping performance comparable to full precision Adam.
  • For deep learning practitioners: Adjust learning rates when switching between Adam and sign-based optimizers by following the authors’ simple transfer rule.

Authors

Yihe Zhou, Tongtian Zhu, Yingxiao Huo, Satya Prakash Dash, Can Wang, Samuel Kaski, Mingfei Sun

Abstract

Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$β$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$β$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.