Language model parts adjust predictably after removal of signals
Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
Machine LearningComputation and Language
Summary
When a piece of a language AI is removed, other pieces seem to adjust as if fixing the problem. The authors explain that this apparent self-fixing is actually a predictable response already present before removal. They found a simple mathematical rule that describes how other parts increase or decrease their activity when one part is changed. This rule holds true across several AI models and different components inside those models. So, the so-called self-repair is actually just normal balancing behavior in the model's parts.
What this means in practice
- •For machine learning engineers: Predict the downstream effects of removing or modifying model components to better debug or optimize language models.
- •For ai system maintainers: Design more reliable model interventions by understanding inherent balancing responses within transformers to improve model robustness.
Authors
Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu
Abstract
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.