Targeted unlearning improves forgetting in large language models per input
TULIP: Targeted LLM Unlearning at Layers Identified Per-Input
Artificial Intelligence
Summary
Large language models store knowledge across many layers inside them. Current methods that try to make these models forget certain information focus on just one fixed layer, but this might miss where the model actually forms the answer. The authors found that the important layer varies depending on the input. They developed TULIP, a method that finds exactly where to intervene for each input, which unlearns the target information more effectively than existing methods. This approach works well across different models and stays strong against attempts to trick the model with paraphrasing or compressed models.
What this means in practice
- •For ai developers: Remove specific unwanted knowledge from large language models more precisely by targeting different internal layers per input.
- •For machine learning engineers: Use per-input layer identification as a modular component to improve existing methods that modify model knowledge without retraining from scratch.
Authors
Yejin Kim, William F. Shen, Seokwon Jung, Daeun Park, Seong Joon Oh
Abstract
Representation-level unlearning intervenes on the intermediate hidden states of LLMs. Although knowledge is distributed across layers, existing methods operate at a single fixed layer for the entire forget set. We ask whether such a fixed layer is sufficient. To answer this, we design a hijacking experiment that grafts hidden states of the target model into an oracle trained only on the retain set. The oracle cannot produce the forget answer on its own, yet it produces the answer from the grafted state. Thus, the answer is formed at an intermediate layer and merely read out afterward, so unlearning should focus on formation, not readout. Moreover, the layer where formation ends varies widely across inputs. Motivated by these findings, we propose Targeted Unlearning at Layers Identified Per-input (TULIP). For each input, TULIP uses the logit lens to locate the formation-readout boundary and removes the hidden state's alignment with the forget answer's unembedding vector there. TULIP consistently outperforms output- and representation-level baselines on TOFU, PISTOL, and WMDP across Llama, Qwen, and Zephyr models. It also remains robust to paraphrase and quantization attacks. Beyond standalone use, its per-input layer selection serves as a plug-and-play component that further improves existing methods.