Massive activations in transformers persist due to read blind spots
Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers
Computation and LanguageMachine Learning
Summary
Some features in transformer AI models stay very active across many layers even though the model can turn them off. The authors found that parts of the model ignore these features when reading data but still write or add to them, creating an imbalance that lets these activations keep building up. This behavior starts early in training and is maintained on purpose by the model. Changing where this ignoring happens causes the model to adjust elsewhere, but these big activations remain.
What this means in practice
- •For machine learning engineers: Improve transformer model debugging by targeting read-write asymmetries that cause persistent feature activations.
- •For ai safety teams: Guide attempts to control internal feature growth in large models by understanding how read-blindness prevents correction.
Authors
Swagatam Mukhopadhyay, Vishal Vivek Saley, Vraj Parikh, Mausam
Abstract
Massive activation features (MAs) in Transformers are extreme-value residual-stream features that persist across layers despite the model's ability to suppress them. Why do they survive? Our investigation using an operator-level mechanistic analysis of attention and feed-forward (FFN) blocks reveals that these blocks systematically ignore MA coordinates while reading, but not while writing; creating a read-write asymmetry that blocks corrective feedback while allowing continued accumulation. We find that both attention and feed-forward layers have this read-blindness, and contribute to the emergence and persistence of MAs. To validate prior work that hypothesized that FFN's amplification abilities is the primary reason for MAs (Sun et al., 2026), we analyze the model checkpoints during learning. Contrary to our expectation, read-blindness emerges before FFN amplification, suggesting that it acts upstream in the MA mechanism. We further contribute gradient analysis to link this behavior to surprising asymmetries in the loss landscape, concluding that the model actively maintains this read-blindness. Finally, we find that removing read-blocking at different locations induces compensatory shifts elsewhere, but MAs still persist.