Language should be separate from thinking in multimodal AI models
Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition
Computation and Language
Summary
Understanding how language fits in AI systems that process multiple types of information is tricky. The authors study how human language affects perception and thought to draw lessons for AI. They find that language acts like a shared codebook, helping compress information but not replacing deeper thought or sensory processing. Their experiments show that AI models handle conflicting visual and language cues differently, with some relying too much on language. They suggest that language in AI should be used mainly at the model's boundaries and in shared communication, not as the core way the model internally thinks.
multimodal modelslanguage modelsperceptioncognitioncodebookcompressioncue conflictideal observertokenshared representation
Authors
Peng Xie
Abstract
Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should keep all of these positions depends on what language does to the system that uses it. The one system with a century of data on that question is the human. We review what language does to human perception, the brain, and thought, and read the same evidence against multimodal models and language models. Throughout, we treat language as a compressor that runs on a shared codebook: a word is an index, the content is in the receiver, and a community maintains the codebook. In humans the compression is measurable, learning the codebook reorganizes the senses, and thought survives the loss of language. We then measure the rule that models apply when two cues disagree, with cue-conflict experiments on six vision-language models and two robot policies. Surviving cues are weighted in the order their reliabilities prescribe, at 11 to 82\% of the ideal observer's slope, and many answers copy the text. One policy family drops a cue that adds no information beyond the others rather than down-weighting it, another keeps it at a weight that fails when the cues conflict, and a visual cue that identifies the task in every training frame is never learned, because the language pathway already fits the data. Language models are the best current models of the human language network, and they have entered the human speech community, shifting word frequencies while alignment narrows their conceptual diversity. We close with seven implications for token-based systems. Language belongs at a model's boundary and in the shared codebook, as in the brain, not as its internal representation; the price of leaving the codebook inside is auditability.