Papers for

multilingual speech service developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Audio deepfake detection improves across languages using language orthogonalization

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

Abstract: Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.

Tue 15 SeptComputation and Language
The gist
Detecting fake audio is hard, especially when the audio is in languages that the system wasn’t trained on. The authors found that current speech models mix language-specific patterns with signs of fakeness, confusing the system. They developed a way to separate the language part from the fake detection part, improving detection for languages not seen during training. This method works better when the new language is very different from those in the training set.
Open → 2609.16458v1