Audio language models get smarter at refusing harmful commands

AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks

SoundCryptography and Security

Summary

Large audio-language models can understand and process spoken content but sometimes fail to reject harmful or unsafe audio commands, a problem called jailbreaks. The authors found these models recognize risky inputs inside but fail to act on that knowledge to refuse unsafe requests. They created AEGIS, a method that detects risky audio in the middle layers of the model and then triggers safety measures to stop undesired answers. This approach greatly reduced unsafe responses across multiple models and test sets without blocking many harmless requests.

What this means in practice

  • For speech system developers: Enhance audio-based AI assistants to better detect and refuse harmful or inappropriate audio inputs by selectively activating safety controls internally.
  • For security teams at ai providers: Deploy AEGIS to decrease the chance of audio-language models producing unsafe outputs when facing adversarial or jailbreak audio inputs.

Authors

Yu-Ling Liao, Tzu-Chin Chiu, Zong-You Chen, Chi-Lei Tsai, Shao-Yuan Lo

Abstract

Large audio-language models (LALMs) expand language models to process and interpret audio, but also expose them to heterogeneous audio jailbreaks. We ask whether successful jailbreaks reflect failures to recognize harmful intent or failures occurring after such recognition. Layer-wise probing reveals the latter: risk-related information remains decodable from intermediate representations, yet the internal risk signal fails to translate into refusal in later-layer processing. We identify this discrepancy as the risk-to-refusal gap. Building on this finding, we propose AEGIS, a detect-then-intervene defense whose mid-layer risk gate selectively activates downstream safety adapters. Across six LALMs and three heterogeneous audio jailbreak benchmarks, AEGIS reduces the average unsafe rate from 17.9% to 0.4%, while causing only a marginal increase in over-refusal on benign inputs. These results establish selective internal intervention as an effective path toward more robust refusal in LALMs. The code is available at https://github.com/azzzzliao/aegis-audio-defense.