Boundary and intra-segment learning improves partial audio deepfake detection
Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization
Sound
Summary
Audio deepfakes can be tricky to spot when only parts of the speech are faked. The authors introduce a new method called BISL that looks not just at the boundary between real and fake speech but also studies the entire authentic and fake segments. This helps the system better tell where the fake parts are, even if changes are subtle. Their approach shows better accuracy than earlier methods on several tests.
What this means in practice
- •For speech security teams: Detect and localize partially tampered audio in communication systems to improve security and trust.
- •For forensic audio analysts: Identify and mark the exact manipulated segments within audio recordings for legal and investigative purposes.
Authors
Zhe Ye, Xiangui Kang, Minhua Huang, Kai Wu, Kong Aik Lee, Chng Eng Siong
Abstract
Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.