Papers for
deep learning practitioners
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Adam optimizer behavior revealed through a stable ratio structure
The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
Abstract: Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$β$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$β$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
Data-driven tuning method improves neural network learning performance
A Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks
Abstract: We study data-driven early stopping for spectral regularisation methods in the classical non-parametric regression setting. Building on the discrepancy principle, we propose a multi-scale stopping rule that applies to general kernel estimators and show that, unlike previous approaches, it achieves full adaptivity over all smoothness levels in the well-specified case. A key contribution of our work is an extension based on random feature approximations, which reduces computational cost on large datasets while preserving minimax-optimal statistical guarantees. Our procedure not only selects an optimal stopping time but also provides a fully data-driven choice of the number of random features needed to achieve optimal rates. Through the established connection between random features and neural networks in the neural tangent kernel regime, our method further yields a principled, data-driven recommendation for the network width. We prove that the resulting simultaneously chosen width and stopping time allow neural networks to attain minimax-optimal learning rates without prior knowledge of smoothness or capacity parameters.