Emergence Invariance: From Symbolized Thought to Interface Refinement

2026-08-03Artificial Intelligence

Artificial IntelligenceMachine Learning
AI summary

The authors explain that language models, like large AI systems, generate thought-like reasoning based on a limited set of rules and inputs, which can miss some human thinking details. They explore how improving these models by scaling up can reduce some gaps but cannot fully replace missing foundational information. Their theory, called the Symbolization--Substructure Thesis, shows why some knowledge limits cannot be overcome by just making models bigger. Experiments confirm that adding key distinctions or better memory helps performance a lot, while just more scale alone has limits. This work connects many ideas about how AI reasoning and memory work together.

Large language modelsEmergenceSymbolization--Substructure ThesisIn-context learningCompensation gapInformation theoryMemory in AIBayesian inheritanceReasoning controlCognitive architecture
Authors
Yi Liu
Abstract
Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition. Large language models trained at scale exhibit compensatory emergence: sparse architectural primitives support in-context learning, multi-step reasoning, tool use, and chain of thought. Yet a language-first probabilistic architecture inherits substantive, substrate, and high-level incompletenesses relative to human cognition. Their coexistence makes an LLM a human-like thought-form generator that reconstructs increasingly human-like reasoning forms from an incomplete substrate. We ask whether emergence can compensate for every missing distinction. We formalize the philosophical premise as the Symbolization--Substructure Thesis and introduce emergence invariance. For a scale-indexed family acting through a shared task interface $φ$, $\mathcal{R}_s^*=\mathcal{R}_φ^*+C_s$: scale can reduce the compensation gap $C_s$, while a positive interface floor $\mathcal{R}_φ^*$ persists. We prove that, under a fixed input law, one interface is universally no less informative exactly when its completed information $σ$-field refines the other, and that total compensation occurs exactly when both the interface floor and asymptotic compensation gap vanish. The framework unifies existing results on grounding, memory, position, attention, Bayesian inheritance, scientific abduction, and reasoning control. In a matched DeepSeek V4-Flash API study, thinking improves pointer chasing from $0/16$ to $14/16$ when relevant distinctions are available; exact observational twins remain at their $50\%$ construction floor; and restoring decisive memory moves matched performance from $50\%$ to $100\%$. These results provide initial evidence for the predicted separation between scaling within an interface and refining the interface itself.