AIM-ZO improves efficiency of tuning large language models without backpropagation

AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning

Machine Learning

Summary

Fine-tuning huge AI language models usually needs a lot of memory and complicated calculations. The researchers found a way to adjust these models using only forward checks, skipping the usual backpropagation step. They do this by focusing updates in helpful directions informed by the model’s activations during use. This approach, called AIM-ZO, uses these signals to keep track of useful update paths and tests changes efficiently. Their tests show better performance on multiple models and language tasks compared to previous methods.

What this means in practice

  • For machine learning engineers: Fine-tune very large language models with lower memory use by applying activation-informed low-dimensional updates without backpropagation.
  • For cloud ai service providers: Deploy efficient large language model fine-tuning that saves server resources by limiting costly backpropagation and memory storage.

Authors

Yue Xie, Zhi Zheng, Yunpeng Ba, Xuyang Wu, Xialiang Tong, Zhichao Lu, Tao Zhong, Zhenkun Wang

Abstract

Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly informative directions. To make these evaluations more informative, existing ZO methods restrict perturbations to low-dimensional subspaces. Yet the quality of these subspaces is critical: overly compressed or poorly maintained spaces can miss useful update directions. To obtain a high-quality subspace for ZO updates, this paper proposes AIM-ZO, a ZO fine-tuning method based on Activation-Informed Subspace Maintenance. AIM-ZO uses forward activations as local directional information and continuously integrates them into a broad, evolving subspace over training. To access broader gradient-relevant structure while keeping individual perturbations low-dimensional, AIM-ZO activates only a smaller set of shared and sampled directions, decoupling the maintained width from the active width. We evaluate AIM-ZO across 5 LLMs and 11 downstream tasks under matched forward-evaluation budgets; its six-task average exceeds the strongest fully evaluated ZO baseline by 1.26 percentage points on OPT-2.7B and MeZO by 2.85 percentage points on OPT-30B. Our code is available at https://github.com/EkkoXy/AIM-ZO