SoftServe improves deep learning training with scalable quasi-Newton methods

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Machine LearningArtificial Intelligence

Summary

Optimizing deep neural networks can be very hard because they have many parameters and their math is complicated. The authors created SoftServe, a new way to improve training by estimating curvature to guide updates better, even when the usual assumptions don’t hold. It uses efficient math tricks that work well with GPUs and can handle very large networks. This approach often achieves better performance than popular existing methods on difficult learning tasks.

What this means in practice

  • For deep learning engineers: Train large and complex neural networks more effectively by using SoftServe’s scalable quasi-Newton optimization that handles difficult curvature without extra tuning.
  • For machine learning platform developers: Integrate SoftServe’s GPU-friendly matrix operations to accelerate training workflows for ill-conditioned models at large scale.

Authors

Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower

Abstract

Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.