Audio deepfake detection improves across languages using language orthogonalization

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

Computation and Language

Summary

Detecting fake audio is hard, especially when the audio is in languages that the system wasn’t trained on. The authors found that current speech models mix language-specific patterns with signs of fakeness, confusing the system. They developed a way to separate the language part from the fake detection part, improving detection for languages not seen during training. This method works better when the new language is very different from those in the training set.

What this means in practice

  • For voice security teams: Improve audio deepfake detection tools to better recognize fake speech in languages not covered by existing training data.$Commercial implications: Enables development of more reliable multilingual anti-spoofing products for security and authentication providers.
  • For multilingual speech service developers: Enhance voice-based services by reducing false alarms when detecting audio manipulation across different languages.

Authors

Minu Kim, Ji Sub Um, Hoirin Kim

Abstract

Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.