Extender reduces memory needs for transformer attention with new channel

The Extender: A Log-Structured Transformer

Machine LearningArtificial Intelligence

Summary

Transformers are a type of AI model that process information in layers, usually passing one big summary through all layers. The Extender adds a second, smaller channel that keeps track of extra details from each layer. This makes it possible to use much less memory when looking at long sequences of information. The authors found that the Extender works as well as, or better than, standard Transformers especially on tasks needing longer context, while drastically cutting memory use.

What this means in practice

Authors

Jakob Eriksson

Abstract

We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each layer $\ell$ emits both a residual update $δ_\ell$ which is added to $\mathbf{h}$, and a much smaller extension $ε_\ell$ which is appended to $\mathbf{x}$. While both the FFN and $\mathbf{q}$ see $\mathbf{h}$, the attention $\mathbf{kv}$ projections take only $\mathbf{x}$ as input. As a result, the fully extended $\mathbf{x}$ contains the complete input for the $\mathbf{kv}$ projections of all layers, reducing the persistent attention memory footprint from $2Ld_{model}$ to $\sum|ε_\ell|$. We find that with $|ε_\ell|=32$, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is $104\times$ smaller than MHA. The memory savings grow with model width.