Auditable records improve trust in speech deepfake detection scores
From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection
SoundComputation and Language
Summary
Speech deepfakes are fake audio clips that sound like real people, making it hard to tell if they are real or fake. The authors show that instead of giving just one score for each audio clip, detectors can provide a detailed record that explains why a clip is scored a certain way. This record includes different types of evidence and helps reviewers understand which clips need more attention. Their approach improves detection accuracy and helps catch more errors while still giving an easy-to-use final score.
speech deepfakedeepfake detectionscoring systemauditable decisionretrieval supportcalibrationerror ratespeaker profileASVspoof dataset
Authors
Mengzhe Geng, Yujia Lu, Patrick Littell, Manuela Kunz, Xie Chen
Abstract
Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress in speech deepfake detection, but most detectors still end with one score per utterance. That score is useful for ranking systems, yet it says little about why a borderline item should be trusted, deferred, or reviewed. Two utterances can fall in the same score band for different reasons, for example because passive and retrieval evidence disagree or because the keyed probe is unavailable. We ask whether the final decision can remain scalar without discarding that provenance. We answer this question with an auditable decision record that carries four aligned cues into a late calibration step: a passive detector score, a conditional keyed-probe score on a marked derivative, retrieval support, and a speaker-profile margin, together with explicit disagreement coordinates. On the 4,080-example ASVspoof 5 Track 1 matched subset, the fixed retrieval-augmented rule improves on retrieval-only evidence, from 15.84 percent to 11.91 percent EER, and late calibration over the full record reaches 8.43 percent EER. At a 33.75 percent review budget, the exposed cue union covers 82.85 percent of the calibrated model's errors. The best passive WavLM run still reaches 6.71 percent EER, so we do not present the decision record as a stronger standalone detector. Its contribution is to preserve the evidence behind each surfaced utterance while still producing one operating score for thresholding and review.