Graph-based wavelet tool cuts parameters for speech deepfake detection

WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection

Artificial IntelligenceSound

Summary

Detecting fake speech can be tricky because the part of the system analyzing sounds needs to focus on useful details. The authors created a new way to organize wavelet scattering data, keeping the connections between sound components in a graph structure. This lets their method capture important speech clues using fewer adjustable parts while still working well, especially when tested on unfamiliar examples. Their approach helps computers better spot faked voices by keeping the sound data’s natural organization.

What this means in practice

  • For voice security teams: Implement a more compact and effective front-end module to improve detection of manipulated speech audio using fewer computational resources.
  • For audio software developers: Integrate a topology-preserving wavelet scattering graph interface for building robust speech deepfake detection features in audio analysis tools.

Authors

Kwok-Ho Ng, Tingting Song, Bingwen Feng, Zhihua Xia

Abstract

The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.