Shared-output local learning speeds billion-parameter model training
SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
Machine LearningDistributed, Parallel, and Cluster Computing
Summary
Training very large language models usually requires passing signals backward through all layers, which can slow down training and use lots of memory. The authors found that training each part independently without sharing information doesn't scale well for very big models. They introduced a new method called SOLO that shares output information without needing full backward passes, making training faster and more memory efficient. Their approach performs nearly as well as traditional methods on large models and data.
What this means in practice
- •For machine learning engineers: Speed up training of large language models by reducing memory needed for backward passes using SOLO's shared-output local learning.
- •For data center operators: Increase throughput during language model pretraining by enabling larger micro-batches through reduced memory locking in pipeline stages.
Authors
Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li
Abstract
Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict the target through its own readout, but has not scaled to billion-parameter pretraining. We identify these private readouts as a key weakness, since they leave each module without information from deeper modules. We propose Shared-Output LOcal learning (SOLO), which replaces them with a shared, read-only copy of the final module's readout, the only one trained on the output of the whole network. Taken from the previous step, the copy transmits information from the final module without passing gradients between modules or reintroducing update locking. SOLO approaches backpropagation on Transformers of 340M to 2B parameters pretrained on 15B tokens, staying within one point in average zero-shot accuracy with a perplexity gap that narrows with scale. Readout ablations attribute SOLO's improvement over private readouts to sharing. Without update locking, each of p pipeline stages holds activations for O(1) micro-batches instead of O(p). The freed memory permits larger micro-batches, which reach up to 1.44x the best measured throughput of pipeline backpropagation on the same partition. To our knowledge, SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining. Local learning thus becomes a practical alternative to backpropagation for large-scale pretraining.