Papers for

voice security teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Speech deepfake detector improves spotting fake voices on new sources

GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

Abstract: Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.

Mon 28 SeptSoundArtificial Intelligence
The gist
Realistic fake voices made by AI can trick people, so detecting them is important. The authors found that existing detectors struggle with voices from new sources and miss subtle fake signals. They designed a new system called GLAD that looks closely at both overall speech patterns and tiny details to better spot fakes. GLAD also adapts to different kinds of voice recordings and uses data tricks to avoid being fooled by background noise. Their tests show GLAD works better than current methods, especially with unseen types of fake audio.
Open → 2609.35411v1

Speech deepfake detection improves by fixing training conflicts

DGS-MLDG: Domain Gradient Surgery Guided Meta-Learning for Domain Generalization in Speech Deepfake Detection

Abstract: Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through an asymmetric projection strategy. DGS removes the destructive component from the meta-test gradient, ensuring a conflict-free optimization trajectory versus the meta-train gradient. Furthermore, we introduce layer-wise DGS (LW-DGS), an efficient variant of DGS that dynamically identifies and intervenes only conflict-prone layers. Extensive experiments on challenging benchmarks demonstrate that DGS-MLDG and LW-DGS-MLDG achieve an average relative EER reduction of 5.29% and 4.04%, respectively.

Sun 27 SeptSound
The gist
Detecting fake speech can be tricky because the way people speak changes in different situations. The authors found that one popular way to train detectors runs into problems because two parts of the training disagree and slow progress. They created a new method that stops these disagreements from messing up training. This helps the system get better at spotting fake speech even when the speaking style changes.
Open → 2609.33706v1

Speech deepfake detection improves using combined acoustic features

What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

Abstract: The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.

Sun 27 SeptSoundArtificial IntelligenceMachine Learning
The gist
Detecting fake speech is getting harder because new ways to create it sound more realistic. The authors studied many sound features to find which ones best spot fakes across different speech generation methods. They found that a mix of two types of features, including one called no-vocals residuals, works especially well when the fake speech method is new or unseen. Combining these features into a single detector greatly reduced errors in spotting fake speech. This approach could help build stronger protections against audio deepfakes.
Open → 2609.33375v1

Graph-based wavelet tool cuts parameters for speech deepfake detection

WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

Abstract: The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.

Thu 24 SeptArtificial IntelligenceSound
The gist
Detecting fake speech can be tricky because the part of the system analyzing sounds needs to focus on useful details. The authors created a new way to organize wavelet scattering data, keeping the connections between sound components in a graph structure. This lets their method capture important speech clues using fewer adjustable parts while still working well, especially when tested on unfamiliar examples. Their approach helps computers better spot faked voices by keeping the sound data’s natural organization.
Open → 2609.29372v1

Audio deepfake detection improves across languages using language orthogonalization

Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

Abstract: Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.

Tue 15 SeptComputation and Language
The gist
Detecting fake audio is hard, especially when the audio is in languages that the system wasn’t trained on. The authors found that current speech models mix language-specific patterns with signs of fakeness, confusing the system. They developed a way to separate the language part from the fake detection part, improving detection for languages not seen during training. This method works better when the new language is very different from those in the training set.
Open → 2609.16458v1