Machine unlearning aims to remove specific influences in large language models
Machine Unlearning for Large Language Models: Foundations, Advances, and Agentic Extensions
Cryptography and Security
Summary
Sometimes developers want to make large language models forget specific information without hurting the model's other skills. This paper looks at different ways to do this, compares tools and tests, and suggests a framework to analyze how well unlearning works. The authors find current tests are not enough to prove that the unwanted data is truly removed, especially when models update or work with extra tools. They suggest better evaluation methods to track if the forgotten information can come back.
What this means in practice
- •For machine learning teams: Check and improve procedures that remove data influence from deployed language models without breaking other capabilities.
- •For software engineers: Design systems that interact with language models using tools or memory and ensure targeted forgetting works reliably across updates.
A survey. It maps existing work.
Authors
Xiaoyu Xu, Minxin Du, Li Bai, Junxu Liu, Yaxin Xiao, Kun Fang, Liu Yang, Huadi Zheng, Peizhao Hu, Qingqing Ye, Haibo Hu
Abstract
Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target locations, interventions, and supported claims. A seven-stage lifecycle and six evidence dimensions guide comparison. The review shows that target construction, retained data, and recovery tests affect reported outcomes. Evidence from model evaluations remains insufficient to establish removal across external state and subsequent updates, motivating evaluation that tracks dependencies and tests whether target influence returns.