Papers for

audio forensic analysts

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Speech deepfake detection improves by fixing training conflicts

DGS-MLDG: Domain Gradient Surgery Guided Meta-Learning for Domain Generalization in Speech Deepfake Detection

Abstract: Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through an asymmetric projection strategy. DGS removes the destructive component from the meta-test gradient, ensuring a conflict-free optimization trajectory versus the meta-train gradient. Furthermore, we introduce layer-wise DGS (LW-DGS), an efficient variant of DGS that dynamically identifies and intervenes only conflict-prone layers. Extensive experiments on challenging benchmarks demonstrate that DGS-MLDG and LW-DGS-MLDG achieve an average relative EER reduction of 5.29% and 4.04%, respectively.

Sun 27 SeptSound
The gist
Detecting fake speech can be tricky because the way people speak changes in different situations. The authors found that one popular way to train detectors runs into problems because two parts of the training disagree and slow progress. They created a new method that stops these disagreements from messing up training. This helps the system get better at spotting fake speech even when the speaking style changes.
Open → 2609.33706v1

TEMA improves tracking and timing answers in multi-audio dialogs

TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs

Abstract: Multi-turn, multi-audio temporal question answering requires models to track target events across follow-up questions, recording switches, and historical references, recovering complete instances and their boundaries for temporal calculation and comparison. We propose TEMA, which connects event perception with evidence-based answering through Route, specifying the audio scope, and Span, describing all relevant intervals as conditional audio captions. We construct TEMA-Dialog with 40,704 dialogs and per-turn evidence and answer supervision, and TEMA-Bench for joint evaluation of evidence and final answers. Training combines temporal grounding initialization, full-dialog supervised fine-tuning, and completeness-first Span-only GRPO. Experiments on Qwen2.5-Omni and AF-Next show improved temporal question answering, particularly event localization and cross-audio comparison. Reinforcement learning applied solely to evidence further improves interval recovery and answer accuracy.

Thu 24 SeptSound
The gist
Answering questions about events happening in multiple audio clips over several turns is tough because you have to remember what happened before and when. The authors created TEMA, a new system that links hearing events with clear evidence to answer these tricky questions. They built a large dataset and tests to train and check TEMA, which helps spot events and compare them across different audio sources more accurately. They also improved it by using a special training method that focuses on finding evidence first.
Open → 2609.30029v1