Language models recover meaning from corrupted inputs spontaneously
Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs
Computation and LanguageArtificial Intelligence
Summary
Sometimes language models give the right answer even if their input has mistakes like missing words or typos. The authors studied how this happens inside the model and found a two-step process where early parts spot errors and later parts fix them. This repair ability appears naturally, even wenn the models were only trained on perfect, clean text. They also made tools that can predict when the model will fail on bad inputs, which could help avoid wasted work.
What this means in practice
- •For llm deployment teams: Predict failure of language models on corrupted inputs to reduce unnecessary checking or compute.
- •For ai model trainers: Improve robustness of language models to input errors by fine-tuning with moderate corruption to increase tolerance and linearity.
Authors
Pranjal Garg, Jacob Beck
Abstract
Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges spontaneously despite training exclusively on clean sequences, without corruption training or an explicit denoising objective. We find that context restoration follows a two-phase process: early layers localize effects associated with repair at corrupted positions, while later layers accumulate these effects at uncorrupted positions through the residual stream and ultimately concentrate them at the output position. Repair outcome is predictable from hidden states: cosine alignment with the clean state is highly predictive in attention-only models, while linear probes recover additional information in pretrained LLMs. A linear probe using only the corrupted prompt's first-block hidden state predicts failure with mean ROC-AUC 0.78. This enables failure triage under matched or even partially shifted deployment conditions and may reduce unnecessary verification or computation. Failed examples also show substantially greater nonlinearity along corruption directions. Moderate-corruption finetuning increases corruption tolerance while simultaneously reducing displacement-normalized linearization error, associating improved robustness with a more nearly linear response to corruption.