Extender reduces memory needs for transformer attention with new channel
The Extender: A Log-Structured Transformer
Machine LearningArtificial Intelligence
Summary
Transformers are a type of AI model that process information in layers, usually passing one big summary through all layers. The Extender adds a second, smaller channel that keeps track of extra details from each layer. This makes it possible to use much less memory when looking at long sequences of information. The authors found that the Extender works as well as, or better than, standard Transformers especially on tasks needing longer context, while drastically cutting memory use.
What this means in practice
- •For machine learning engineers: Build transformer models that use significantly less memory when processing long sequences without losing accuracy.
- •For data center infrastructure teams: Reduce memory costs and hardware requirements when deploying large transformer models for tasks involving extensive contextual data.
Authors
Jakob Eriksson
Abstract
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each layer $\ell$ emits both a residual update $δ_\ell$ which is added to $\mathbf{h}$, and a much smaller extension $ε_\ell$ which is appended to $\mathbf{x}$. While both the FFN and $\mathbf{q}$ see $\mathbf{h}$, the attention $\mathbf{kv}$ projections take only $\mathbf{x}$ as input. As a result, the fully extended $\mathbf{x}$ contains the complete input for the $\mathbf{kv}$ projections of all layers, reducing the persistent attention memory footprint from $2Ld_{model}$ to $\sum|ε_\ell|$. We find that with $|ε_\ell|=32$, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is $104\times$ smaller than MHA. The memory savings grow with model width.