Neural language models learn meanings from sentence structures not just words

Neural Language Models Learn the Contextual Distributions of Dependency Structures: a statistical learning theory to compositionality

Computation and Language

Summary

Language models usually learn by looking at word statistics, but it’s unclear how they understand sentence structures that carry meaning beyond words. The authors show that these models can learn patterns of how grammatical structures appear in different contexts, even when those patterns can't be guessed by looking at individual words. They trained models on a made-up language where each structure had unique context patterns and found that models first learn the grammar relationships and then the meanings from context. This suggests that models build understanding by treating grammatical structures as new units for learning meaning.

What this means in practice

  • For nlp engineers: Improve language model training by incorporating learning of grammatical structure contexts to better capture sentence-level meaning beyond word statistics.
  • For software linguists: Design language tools that analyze syntax and semantics by tracking learned dependency structures as new units, enhancing understanding of complex sentences.

Authors

Wang Bojun, Junjie Chen, Holly Jenkins, Elizabeth Wonnacott

Abstract

It is unclear how Neural Language Models (NLMs) acquire the structural meaning encoded by grammatical structures that is independent of lexical semantics. We propose a statistical learning process in which learned dependency structures themselves become new distributional units for subsequent statistical learning. Under this account, once a dependency structure is acquired, the model tracks its contextual distributions. These contextual features reflect the semantic properties of a composite structure. To test this hypothesis, we design a synthetic grammar in which each grammatical structure has distinct contextual distributions that cannot be recovered from the distributional statistics of their component tokens alone. We train a series of BERT-style masked language models on this grammar and examine their developmental trajectory. The results show that models can successfully learn the contextual distributions of composite dependency structures even though they cannot be inferred from token statistics alone. Developmental analysis further reveals a clear developmental trajectory. The learning of the dependency relations that define a grammatical structure consistently precedes the learning of its contextual features. These findings suggest that statistical learning in NLMs is not merely the accumulation of token co-occurrence statistics, but a process in which learned dependency structures become new units of distributional learning. We argue that this process provides a statistical-learning account of how NLMs solve the compositionality problem in language. Finally, we discuss the possibility that this statistical learning process provides an explanatory theory on how language cognition could emerge from pure distributional statistics.