Papers for

forensic audio analysts

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Speech deepfake detector improves spotting fake voices on new sources

GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

Abstract: Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.

Mon 28 SeptSoundArtificial Intelligence
The gist
Realistic fake voices made by AI can trick people, so detecting them is important. The authors found that existing detectors struggle with voices from new sources and miss subtle fake signals. They designed a new system called GLAD that looks closely at both overall speech patterns and tiny details to better spot fakes. GLAD also adapts to different kinds of voice recordings and uses data tricks to avoid being fooled by background noise. Their tests show GLAD works better than current methods, especially with unseen types of fake audio.
Open → 2609.35411v1

Speaker verification system explains voice decisions with acoustic details

CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification

Abstract: Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and supplying explicit acoustic measurements, CoLMbo-SV makes voice comparisons inspectable without restricting verification to the evidence verbalized in its reports. We additionally introduce \textbf{VoxReason}, paired recordings with measured acoustic properties and comparison reports filtered through numerical and qualitative checks, providing supervision for this combined capability. We also develop an evaluation framework that separates what acoustic information a speaker representation encodes, what influences the verification score, and what the generated report discusses. On VoxCeleb1-O, CoLMbo-SV achieves 0.99\% EER, reducing verification error by approximately 80\% relative to the strongest audio-language baseline fine-tuned on VoxReason, while attaining a numerical-grounding score of 0.82. Our analysis further demonstrates that acoustic correctness and decision relevance are distinct properties of an explanation, exposing a gap that numerical-grounding metrics miss. Together, these contributions substantially advance audio-language speaker verification, bring its accuracy toward that of dedicated speaker encoders while adding checkable acoustic reporting, and establish an empirical framework for connecting natural-language explanations to the decisions they explain.

Sun 27 SeptComputation and LanguageSound
The gist
Current speaker verification systems can tell if two voices are from the same person but do not explain how they made that decision. The authors created CoLMbo-SV, which combines a voice recognition model with a language model that produces detailed, understandable reports that refer to specific acoustic measurements. These reports help users inspect and understand the evidence behind verification results without limiting the system’s accuracy. They also introduced VoxReason, a dataset with paired voice recordings, measurements, and verified comparison reports to train and evaluate such models.
Open → 2609.33212v1

Boundary and intra-segment learning improves partial audio deepfake detection

Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Abstract: Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.

Tue 22 SeptSound
The gist
Audio deepfakes can be tricky to spot when only parts of the speech are faked. The authors introduce a new method called BISL that looks not just at the boundary between real and fake speech but also studies the entire authentic and fake segments. This helps the system better tell where the fake parts are, even if changes are subtle. Their approach shows better accuracy than earlier methods on several tests.
Open → 2609.25822v1

Transformer model improves verifying family relations from speech

A New Transformer-Based Approach for Audio-Based Kinship Verification and a New Uncontrolled Mandarin Kinship Speech Dataset

Abstract: Kinship verification is a task involving determining whether two individuals share a first-order kin relation. To tackle this task, we propose CONVTRAP-TN, a new architecture for audio-based kinship verification, and conduct an ablation study on the proposed model. To the best of our knowledge, we are the first to apply the successful transformer architecture to the task of audio-based kinship verification. Furthermore, we also collect a custom speech dataset, ARKIN, which accurately reflects everyday recording conditions. We do this because only a few speech datasets with kinship labels currently exist, all of which either source extremely noisy in-the-wild data from the internet, or instruct speakers to record in specific environments. These settings fail to reflect real-world scenarios where users record on personal devices under unrestrained conditions. Additionally, we perform a series of preliminary baseline experiments on the collected dataset, including speaker verification and recognition, speech recognition, age estimation, and kinship verification, as well as cross-dataset kinship verification experiments to show that existing methods are not robust across datasets.

Sat 12 SeptSoundArtificial IntelligenceMachine Learning
The gist
Determining if two people are closely related just by listening to their voices is a tricky task. The authors created a new method using a transformer-based model called CONVTRAP-TN to better recognize family connections from audio. They also collected a new Mandarin speech dataset recorded in everyday situations to better train and test such systems. Their experiments show current methods struggle when tested on different datasets, suggesting their approach and data reflect real-world use better.
Open → 2609.14145v1