Informational Antilocality and the Locality Bias in LLMs

Computation and Language

Summary

The authors studied how well transformer-based language models can learn special languages called k-antilocal languages, where there is no connection between any sequence of k symbols. They created these languages with different values of k and found that while the models eventually learn them equally well, it takes longer to learn languages with higher k. This suggests that learning relationships between far-apart parts in language is harder, but the models still succeed in the end.

transformerlanguage modelk-antilocal languagesmutual informationcross-entropy lossnon-local dependencieslearning speedlearning convergence

Authors

Andrew McInnerney, Shane Storks, Steven Abney, Richard L. Lewis

Abstract

We consider the ability of transformer-based language models (LLMs) to learn what we call k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols. We construct such languages with increasing $k$, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages. Our findings support the idea that non-local dependencies are more difficult to learn, but the evidence for this bias comes from learning speed rather than learning success.