Speech deepfake detection improves by fixing training conflicts

DGS-MLDG: Domain Gradient Surgery Guided Meta-Learning for Domain Generalization in Speech Deepfake Detection

Sound

Summary

Detecting fake speech can be tricky because the way people speak changes in different situations. The authors found that one popular way to train detectors runs into problems because two parts of the training disagree and slow progress. They created a new method that stops these disagreements from messing up training. This helps the system get better at spotting fake speech even when the speaking style changes.

What this means in practice

  • For voice security teams: Improve fake speech detection systems in security settings where audio varies widely across sources.
  • For audio forensic analysts: Enhance evaluation tools for speech authenticity when audio comes from diverse environments and recording conditions.

Authors

Siqing Qin, Kong Aik Lee, Youzhi Tu, Eng Siong Chng, Man-Wai Mak

Abstract

Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through an asymmetric projection strategy. DGS removes the destructive component from the meta-test gradient, ensuring a conflict-free optimization trajectory versus the meta-train gradient. Furthermore, we introduce layer-wise DGS (LW-DGS), an efficient variant of DGS that dynamically identifies and intervenes only conflict-prone layers. Extensive experiments on challenging benchmarks demonstrate that DGS-MLDG and LW-DGS-MLDG achieve an average relative EER reduction of 5.29% and 4.04%, respectively.