Queries expand and keys shrink in transformer attention mechanisms
Query Expansion and Key Specialization in Transformer Attention Geometry
Artificial IntelligenceMachine Learning
Summary
In transformer models used for tasks like reading text, two components called queries and keys work together to focus attention on important information. This paper finds that during training, queries tend to spread out in how they represent data, while keys become more focused and narrow. This difference makes the attention more precise and confident. The authors also tested controlling this narrowing effect and found it directly affects how sharply the model pays attention, especially early in training.
What this means in practice
- •For machine learning engineers: Adjust attention key projections during transformer training to control focus sharpness and confidence.
- •For natural language processing developers: Improve transformer-based character-level language models by monitoring query and key dimensional changes to optimize attention.
Authors
Vidit Gupta, Siddhesh Nadkarni, Mihik Chaudhari, Vinaya Sawant, Prachi Tawde
Abstract
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms. The question of whether there is an effect from their functional distinction on their geometric development in training remains unanswered. We investigate the problem through the training of small GPT-like Transformers on character-level WikiText-103 for three different depths (4, 6, and 8 layers), three types of initialization for queries and keys, and four random seeds, resulting in 36 runs and 54 trajectories of average layers across seeds. We track the effective dimensionality of those layers using participation ratios and discover that effective dimension of queries expand while keys shrink, and that $PR_Q - PR_K$ is positive in all trajectories studied. In connection to attention, the shrinking of keys leads to a narrower spectrum of $QK^\top$ and more peaked attention weights. In order to determine if this connection is causal or coincidental, we directly control the spectrum of keys during training across five seeds: restricting it to make it shrink sharpens the attention with high directional confidence, while keeping it constant to the level of initial dispersion makes attention softer. Additional token-level checkpoint analyses show that the monotonic paired-contrast trend is not universal across pretrained families, but survives as an early-training regime that later decays over a full pretraining run, and the link between interaction-rank geometry and attention entropy remains visible in several models.