Optimizer improves low-rank adaptation training efficiency and outcomes

Rotated Manifold Optimization for Low-Rank Adaptation

Machine Learning

Summary

Training large AI models can be slow and tricky, especially when adjusting only a small part of the model, called low-rank adaptation. The authors created a new method that views these small parts as existing on a special shape, or manifold, and then optimizes by rotating and normalizing factors efficiently. This helps training reach better accuracy faster for tasks like supervised learning and reinforcement learning. Their method combines ideas from matrix optimization and geometry to improve how the training updates are made.

What this means in practice

Authors

Yuhui Ding, Javier Zazo, James Hensman

Abstract

We propose a novel optimizer for low-rank adaptation (LoRA) that explicitly incorporates the gauge symmetry of low-rank factorization. Our optimizer extends recent matrix optimizers for full-parameter training to the manifold of fixed-rank matrices by interpreting them as normalization under a rotated basis. We show how rotation and normalization can be integrated with the fixed-rank manifold efficiently. Our optimizer converges faster to lower held-out loss and achieves better or comparable downstream performance on both supervised finetuning and reinforcement learning tasks.