Speech deepfake detection improves using combined acoustic features

What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

SoundArtificial IntelligenceMachine Learning

Summary

Detecting fake speech is getting harder because new ways to create it sound more realistic. The authors studied many sound features to find which ones best spot fakes across different speech generation methods. They found that a mix of two types of features, including one called no-vocals residuals, works especially well when the fake speech method is new or unseen. Combining these features into a single detector greatly reduced errors in spotting fake speech. This approach could help build stronger protections against audio deepfakes.

What this means in practice

  • For voice security teams: Improve detection of audio deepfakes by integrating pooled no-vocals residuals with speech feature models in diverse generation contexts.
  • For call center operators: Enhance fraud prevention by deploying robust fake speech detectors that work well against unseen voice synthesis methods.

Authors

Jiajun Xu, Menglu Li, Xiao-Ping Zhang

Abstract

The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.