Adaptive step size improves large model training stability and speed

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Machine Learning

Summary

Choosing how big a step to take when training large AI models is tricky: too small means slow learning, too big can cause problems. The authors combine two methods to pick the direction using known gradient info and then decide the step size by checking the model’s performance nearby. This helps adjust steps without expensive calculations and keeps training stable and efficient. Their method works well across different models and data compared to always using fixed step sizes.

What this means in practice

  • For machine learning engineers: Improve training efficiency and stability of large neural networks by adapting step sizes using combined zeroth- and first-order optimization methods.
  • For optimization software developers: Incorporate adaptive step-size selection algorithms that use minimal extra evaluations to enhance optimizer performance without full line searches.

Authors

Cristian McGee, El Houcine Bergou, Aritra Dutta

Abstract

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.