Benchmarking unlearning methods for text to image diffusion models
eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models
Machine LearningArtificial IntelligenceComputer Vision and Pattern Recognition
Summary
It can be hard to compare different ways of removing specific ideas or concepts from text-to-image AI models because each method is tested differently. The authors created eval-unlearn, a free tool that tests and compares 12 known unlearning methods using a shared set of tests. This tool checks how well these methods erase concepts, keep image quality, resist attacks, and retain other knowledge. They also set up a public leaderboard to see how these methods perform on removing nudity. This helps developers understand trade-offs and improve unlearning in these AI models.
What this means in practice
- •For ai model engineers: Compare and select effective concept removal methods for text-to-image AI using a unified evaluation framework.
- •For content moderation teams: Assess and improve AI capabilities to remove unwanted or unsafe image concepts such as nudity in generated content.
Authors
Mansi, Nikhil Raghavan, Zixia Huang, Kai Sheng Ong, Ji Shen Lim, Brandon Siao Xiang Ling, Francesco Leofante
Abstract
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models. eval-unlearn integrates twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention, alongside nine complementary evaluation metrics covering erasure efficacy, adversarial robustness, generative quality, and concept retention. Its plugin architecture lets third-party techniques and metrics self-register without modifying the core framework, and its streaming, batched pipeline supports efficient evaluation of both standard NSFW concepts and arbitrary general concepts. As a further contribution, we release a public leaderboard on HuggingFace along with an interactive tool for real-time evaluation of unlearning techniques. The leaderboard compares nudity concept erasure case study across all twelve techniques, exposing significant accuracy-quality trade-offs that are obscured by heterogeneous evaluation. eval-unlearn is released under the MIT license; the package, code, leaderboard, and documentation are all available at https://eval-unlearn.readthedocs.io.