Large language models use attention heads to detect internal changes

A mechanistic study of language model introspection

Computation and LanguageArtificial Intelligence

Summary

Sometimes, large language models notice when their own internal parts change, even if their input stays the same. The authors studied how the models identify and locate these internal changes using a test where they secretly altered states inside the model at different spots. They found two specific groups of attention heads that help the model notice changes and decide exactly where they happened. Understanding these internal mechanisms helps explain how language models can introspect about their own processing.

What this means in practice

Authors

Jiahong Zou, Xiangkun Sun, Lingkai Kong, Tonghan Wang

Abstract

Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.