Language model pruning improves output sensitivity and accuracy

Output-aware Residual Stream Pruning for Large Language Models

Machine Learning

Summary

Shrinking parts of large language models can make them run faster, but not all parts being removed are equally important for the final answer. The authors propose a smarter way to decide what to trim, by looking at how much changes in these parts actually affect the model’s output. Their new method picks which parts to keep based on how sensitive the final results are, not just on preserving internal activity. Tests show this approach helps language models keep better performance and produce more reliable outputs when they are compressed.

What this means in practice

Authors

Chayne Thrash, Kevin Chen, Soheil Kolouri

Abstract

Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.