Hybrid attention layers impact multilingual model learning speed and alignment

Multilinguality in Hybrid Attention LLMs

Computation and LanguageArtificial Intelligence

Summary

Handling very long sentences in multiple languages is tough for AI language models because traditional methods use a lot of computing power. The authors studied models that mix full attention and recurrent attention layers, which changes how these models understand different languages. They found that the order of these layers affects how well the model connects languages and learns, with starting on full attention layers being better. This insight can make training multilingual AI models faster and improve their understanding across languages.

What this means in practice

Authors

Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz, Nanyun Peng

Abstract

In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations of full attention and alternatives based on recurrence. This work presents a first study of how hybrid attention impacts the multilinguality of LLMs. Beyond the impact on long sequences in poorly tokenized languages, our study is motivated by the possibility that the inductive biases of the recurrent state alter linguistic processing. Our interpretability analysis confirms this, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers. Across diverse models, we notably observe a pronounced spike in cross-lingual alignment around the first full-attention layer. These findings lead us to question the conventional ordering of attention layers. In distillation experiments on multilingual data, all alternative layer orderings outperform the standard throughout training, learning up to 2.5X faster. These stark, replicable results prompt our theory that multilingual models would benefit from starting with a full-attention layer rather than recurrent layers.