AEGIS: Awareness-Enhanced Guidance for Iterative Safeguard
2026-07-20 • Computation and Language
Computation and Language
AI summaryⓘ
The authors study a method called AEGIS that helps rewrite harmful text using special hints about which parts are offensive in English, Chinese, and Korean. Instead of claiming the best results, they explore how giving these hints affects the balance between making text less toxic and keeping its original meaning. Their findings show that using these hints can help or hurt depending on the writing model and language used. This means that such detailed control is helpful sometimes but not always.
text detoxificationspan-level rationalemultilingual NLPtoxicity reductionmeaning preservationgenerator backbonestructured guidanceAEGIStext rewritinglanguage models
Authors
Kyungwon Park, Sangmin Lee, Heejae Chon, Hyungu Kang
Abstract
Span-level rationales are often assumed to improve controllability in text detoxification, but it remains unclear when such guidance helps and when it introduces trade-offs. We present Awareness-Enhanced Guidance for Iterative Safeguard (AEGIS) as an exploratory framework for studying span-guided multilingual detoxification across English, Mandarin Chinese, and Korean. AEGIS combines span-level detector outputs with frozen generator backbones, allowing harmful spans, intensity labels, and target attributes to be provided as structured guidance during rewriting. Rather than claiming state-of-the-art detoxification performance, we analyze how span guidance affects the balance between toxicity reduction and meaning preservation across generator families, model scales, and languages. Our results suggest that span-guided detoxification is conditionally useful: explicit rationales change the trade-off between toxicity reduction and meaning preservation, but their effects depend strongly on the generator backbone and the linguistic context. These findings highlight both the promise and the limitations of span-level control signals for multilingual detoxification.