An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures

2026-07-03Machine Learning

Machine Learning
AI summary

The authors studied why deepfake audio detectors, which work well in tests, often fail with real-world audio. They found that models rely too much on non-speech parts of the audio, like background noise, instead of focusing on the actual voice. By carefully changing parts of the audio and analyzing how the models react, they showed that these non-speech sections cause the biggest drops in performance. This helps explain why models struggle when audio conditions change.

deepfake audio detectionacoustic perturbationsnon-speech intervalsshortcut learningdomain shiftdirected graphical modelXLS-R-300MRawGAT-STASVspoof datasets
Authors
Santiago Rubio, Pilar Bello, Dayana Ribas, Antonio Miguel, Eduardo Lleida, Alfonso Ortega
Abstract
While deepfake audio detection systems achieve high performance in controlled benchmarks, their reliability often diminishes in the wild. Prior work shows that dataset-specific artifacts contribute to this gap. Yet, systematic tools to identify which acoustic properties a model exploits as shortcuts remain limited. We propose an intervention-based diagnostic framework, grounded in a directed graphical model, that formally distinguishes confound-driven shortcut dependencies from legitimate domain shift. We operationalise this through controlled acoustic perturbations targeting non-speech structure, spectral content, and signal energy, complemented by corpus-level distributional analysis. Evaluating XLS-R-300M with RawGAT-ST across ASVspoof challenges datasets, we quantify model sensitivity to specific intervention types. Results reveal that non-speech interventions produce the largest performance shifts, confirming non-speech intervals as a dominant shortcut.