Target speech extraction improves with progressive adaptation to real speech
RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing
Sound
Summary
Extracting a person's speech in noisy real-world conversations is hard because training mostly uses fake data that sounds different from real sounds. The authors adapted a method called RemixIT to teach computers progressively using both made-up and real conversation data without perfect labels. Their two-step process first mixes real and synthetic speech information, then refines learning using reliable guesses about real speech signals. This approach made it easier for the system to identify target speech accurately in real conversations, beating older methods in several tests.
What this means in practice
- •For voice assistant developers: Improve extraction of a user's voice from noisy environments for better assistant responses in real conversations.$Commercial implications: Enables voice assistants to work effectively in realistic multi-speaker and noisy settings, improving user experience significantly.
- •For telecommunication engineers: Enhance call quality by extracting target speech from overlapping speakers and background noise in calls.
Authors
Yu Wang, Haixin Guan, Shuang Wei, Yanhua Long
Abstract
Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.