A Formal Limitation on Learning Human Language From Textual Corpora

2026-08-28Computation and Language

Computation and Language
AI summary

The authors studied if it's possible to understand exactly what someone means just by looking at how they say it, without any extra context. They used math from information theory to show that there are limits to how well you can do this, no matter how good your text analysis tools or language models are. Some meaning is impossible to recover without understanding the situation or background, not just the words themselves. Their experiments with different languages and tasks support this idea.

information theoryutterancemeaning recoverylanguage modelsextralinguistic contextzero-pronoun resolutiondiscrete and continuous spacesfeaturizerrepresentation learning
Authors
Emily Cheng, Ryan Cotterell
Abstract
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them; the bounds hold whether the space of meanings is discrete or continuous. Experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference provide empirical evidence in support of the theory.