Adaptive method improves forgetting of sensitive content in large language models

RADNPO: Reference-free Adaptive Negative Preference Optimization for LLM Unlearning

Cryptography and Security

Summary

Large language models sometimes memorize private or sensitive information during training, which can be risky. The authors propose a new technique that not only lowers the chances of the model recalling targeted sensitive information but also carefully redistributes what the model predicts next. This helps avoid repetitive or unhelpful outputs after the model 'forgets' something. Their approach adapts to the model’s confidence in what it originally learned, making unlearning more precise and effective. They tested this method and found it strikes a better balance between removing unwanted knowledge and keeping the model useful.

What this means in practice

  • For ai platform engineers: Improve controls for removing sensitive or proprietary content from language models while preserving model performance after deployment.
  • For content moderation teams: Refine automated methods to ensure models forget specific unwanted information, improving content safety without model degradation.

Authors

Shenghan Tan, Ziyi Zhou, Wenpeng Hu, Mengyuan Zhang

Abstract

Large language models (LLMs) can memorize sensitive, private, or copyrighted content during pre-training, making machine unlearning necessary for removing targeted knowledge. Recent preference optimization (PO)-based unlearning methods improve stability over gradient ascent (GA)-based methods by introducing alignment-style objectives, which effectively suppress the probability of forget targets. However, target suppression alone does not sufficiently constrain the next-token distribution after unlearning. Existing methods provide limited control over how suppressed probability mass is redistributed and insufficiently adapt forgetting strength to target confidence and distributional concentration. Even after target suppression, probability mass may remain concentrated on a few non-target tokens, potentially producing repetitive or uninformative outputs. To address these limitations, we propose Reference-free ADaptive Negative Preference Optimization (RADNPO), which explicitly guides next-token probability redistribution. Specifically, RADNPO contrasts each forget target with alternative tokens favored by the current next-token distribution and adaptively modulates token-level forgetting strength using target confidence and next-token concentration. Experiments on TOFU and MUSE demonstrate that RADNPO achieves a better trade-off between forgetting quality and model utility than current baselines.