Summary
Transformers that use gated linear attention face a challenge because their memory tries to handle both short-term details and long-term meaning at the same time, which limits how well they represent information. The authors developed a new approach called Multi-Scale Gated Linear Attention (MS-GLA), which uses different groups of attention heads to focus on various lengths of text—from small local details to longer dependencies. These groups are combined dynamically to better remember and use information without making the model bigger. Tests show that MS-GLA improves accuracy and understanding on language tasks, especially where recalling details over long text is important.
What this means in practice
- •For language model developers: Build language models that better capture both short-term syntax and long-range context without increasing model size by using multi-scale attention decomposition.
- •For ai system engineers: Improve long-context understanding and recall in AI systems handling extended text inputs by integrating multi-temporal resolution attention mechanisms.
Abstract
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.