Causal routing enables targeted forgetting in large language models

Causal Routing for Unlearning

Machine LearningArtificial Intelligence

Summary

Large language models (LLMs) spread knowledge across millions of parameters, making it hard to remove specific information once learned. The paper introduces Causal Routing for Unlearning (CRU), a way to identify which neurons represent a concept and selectively suppress just those neurons. This targeted suppression allows the model to forget certain content without retraining the entire network or harming other knowledge. CRU uses less memory and better preserves the model’s abilities compared to previous methods.

What this means in practice

  • For ai system engineers: Enable selective removal of specific knowledge from large language models without full retraining, improving efficiency and precision in content control.
  • For privacy compliance teams: Facilitate compliance with data deletion requests by reliably removing targeted information embedded in AI models without degrading overall performance.

Authors

Bardh Prenkaj, Andrea D'Angelo, Davide Mottin, Federico Fontana, Davide Gabrielli, Paola Velardi, Stefano Faralli

Abstract

LLMs cannot forget the way we delete a file. Strangely, we are asked to remove something that was never put anywhere in particular. What the model took from a piece of text is now smeared across billions of weights. Existing methods rewrite all of them to change one thing, and none of them say which part produced that change. To address this, we introduce Causal Routing for Unlearning (CRU) by asking where the concept is expressed in the model and suppressing only that part. One untrained forward pass over the forget set ranks neurons by how their activations vary. Then, small routing modules on those neurons gate and suppress only the concepts that need to be forgotten. In CRU, the base model is frozen, and any change in behavior is caused only by the gated neurons; hence, why the routing is causal. Due to our parameter efficiency (only ~0.01% as many parameters as the base model), unlearning a concept costs 14 GiB, whereas the baselines require 71 GiB. On TOFU, CRU is indistinguishable from the retained model (p > 0.05, KS test) and is never Pareto-dominated, whereas every compared baseline matches its forgetting on the larger-forget batches only by collapsing utility. On RWKU, it achieves an adversarial-probe recall of 0.052, compared to 0.250 for the strongest baseline, meaning the knowledge is gone, not merely harder to reach. Thus, deciding on the intervention at query time, rather than fixing it beforehand, is the axis along which we argue that unlearning should proceed.