Directional study reveals hidden risks in AI model safety training
See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs
Machine LearningComputation and Language
Summary
Large language models trained for safety can still unexpectedly fail when adapting to new topics. The authors studied how these failures develop during training by looking closely at the model’s learning directions and found that certain key words cause sharp risk patterns. They created a method that removes risky learning directions to keep the model safer without losing usefulness. Their tests show this method can reduce unexpected safety failures by a large amount while revealing hidden risk patterns even when the model looks safe.
What this means in practice
- •For llm developers: Reduce unexpected safety failures by eliminating risky training directions in large language models.
- •For ai safety teams: Detect hidden unsafe response patterns in deployed instruction-following models even when behavior looks safe.
Authors
Weiqiao Que, Ruizhe Li, Chengyu Wang, Dakan Wang, Emine Yilmaz, Xiaofeng He
Abstract
Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: https://github.com/WeiqiaoQUE/mechanistic-emergent-misalignment.