Language model outputs traced to key internal components
Matryoshka attribution: Learning to attribute language model outputs to representations and weights
Computation and LanguageMachine Learning
Summary
It is hard to figure out which parts inside large language models actually make their answers. The authors propose a new way called Matryoshka Attribution that learns to find nested groups of important components by minimizing errors. Their method can rank parts by how much they contribute and found small groups that explain model behavior well. They also showed it can identify the weight changes causing an updated model to refuse certain requests, while keeping other abilities intact.
What this means in practice
- •For llm developers: Identify which internal parts or weight changes cause specific behavior changes in large language models.
- •For ml model auditors: Find sparse, transferable circuits responsible for model decisions to improve model transparency and debugging.
Authors
Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts
Abstract
Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.