Speech deepfake detector improves spotting fake voices on new sources

GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

SoundArtificial Intelligence

Summary

Realistic fake voices made by AI can trick people, so detecting them is important. The authors found that existing detectors struggle with voices from new sources and miss subtle fake signals. They designed a new system called GLAD that looks closely at both overall speech patterns and tiny details to better spot fakes. GLAD also adapts to different kinds of voice recordings and uses data tricks to avoid being fooled by background noise. Their tests show GLAD works better than current methods, especially with unseen types of fake audio.

What this means in practice

  • For voice security teams: Protect voice authentication systems by improving detection of AI-generated fake speech across varied and new environments.
  • For forensic audio analysts: Enhance forensic investigations by identifying subtle localized traces of speech forgery that evade typical detection methods.

Authors

Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu

Abstract

Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.