Attention patterns reveal hallucination in large language models
Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
Artificial IntelligenceComputation and LanguageMachine Learning
Summary
Sometimes, large language models make mistakes called hallucinations, where they give wrong or made-up answers. The authors studied how these models share information inside themselves by looking at how their attention mechanisms connect words. They found that hallucinated answers show a breakdown in how the model passes context between words, especially relying too much on certain self-focused attention. This approach helps spot when the model is likely hallucinating by analyzing these connection patterns.
What this means in practice
- •For llm developers: Identify hallucinated outputs in single passes by analyzing attention graph structures during model generation.
- •For machine learning engineers: Incorporate topology-based attention metrics to improve quality control pipelines for language model response validation.
Authors
Amir Jalilifard, Anderson Rocha, Eric Wong, Marcos Medeiros Raimundo
Abstract
In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.