Universal probabilities for sentence structure arise without data
A Data-free Universal Prior over Syntactic Structures
Computation and Language
Summary
This paper shows that some patterns in how sentences are structured can come from how people put words together in real time, rather than from learning specific languages by example. The authors created a model that simulates building sentences word by word without using any actual language data. This model naturally gives higher chances to sentence structures that actually appear in many different languages. It suggests that our brains might have built-in preferences for certain sentence shapes that learning then fine-tunes.
What this means in practice
- •For natural language processing engineers: Initialize probabilistic models of syntax with inherent structure before training on language-specific data to improve parsing accuracy.
- •For cognitive modelers: Incorporate universal syntactic priors into models of language acquisition to better simulate early stages of learning without data exposure.
Authors
Fermín Moscoso del Prado Martín
Abstract
Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing theories estimate the probability of syntactic structures from language-specific data. Whether part of this probability structure can arise independently of language-specific experience remains unknown. Here I show that a universal prior over syntactic structures emerges from a cognitively motivated model of incremental language production, in which words are progressively integrated into syntactic structure through network growth. The resulting prior assigns probabilities to syntactic structures --represented as dependency trees-- without fitting parameters to linguistic data, and assigns higher probabilities to attested than to random trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with probabilities estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific statistical learning. Linguistic experience may therefore refine probabilities that are already structured by the process of language production, rather than create them from an initially uniform space. This identifies a possible cognitive origin for part of the probability distribution over syntactic structures, linking language production and statistical learning while providing a data-independent structural bias for probabilistic models of language.