Large language models use attention heads to detect internal changes
A mechanistic study of language model introspection
Computation and LanguageArtificial Intelligence
Summary
Sometimes, large language models notice when their own internal parts change, even if their input stays the same. The authors studied how the models identify and locate these internal changes using a test where they secretly altered states inside the model at different spots. They found two specific groups of attention heads that help the model notice changes and decide exactly where they happened. Understanding these internal mechanisms helps explain how language models can introspect about their own processing.
What this means in practice
- •For ai system developers: Design AI systems that monitor and respond to internal activation changes for improved model reliability.
- •For natural language processing engineers: Build improved debugging tools that localize internal state changes in language models to help refine model behavior.
Authors
Jiahong Zou, Xiangkun Sun, Lingkai Kong, Tonghan Wang
Abstract
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.