Papers for
telecommunication engineers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Target speech extraction improves with progressive adaptation to real speech
RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing
Abstract: Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.
Polar code decoding improved with offline variance perturbation design
Parallel Successive Cancellation Perturbation-Enhanced Decoding of Polar Codes via Offline Variance Design
Abstract: Successive cancellation perturbation-enhanced (SCP) decoding improves the performance of finite-length polar codes by performing multiple SC decoding attempts with receiver-side perturbations. However, many existing perturbation schemes generate or update subsequent perturbations according to the outcomes of previous decoding attempts, resulting in additional decoding latency. In this paper, we propose an offline variance design (OVD) method for parallel SCP (PSCP) decoding of short- and medium-length polar codes. First, we formulate the exact recovery objective conditioned on ordinary SC failure and classify failed frames by the position of the first genie-aided intrinsic error and the number of subsequent intrinsic errors. We also derive a consistent Gaussian representation of the perturbed channel that preserves min-sum SC hard decisions. Second, we construct a class-based approximation of the recovery objective using Gaussian approximation and backward recursions, accounting for both error correction and new errors introduced by perturbations. We prove that both the exact and analytical objectives are nondecreasing and exhibit diminishing marginal gains as independent branches are added. Third, we develop a greedy algorithm to select variances from a finite candidate set for a given code, signal-to-noise ratio (SNR), and number of perturbation branches. All variances are determined offline, allowing the original SC branch and all perturbation branches to start simultaneously. Simulations for rate-$1/2$ polar codes of lengths $64$, $128$, $256$, and $512$ show that OVD-PSCP achieves lower block error rates (BLERs) than conventional SCP with the same number of perturbation branches. The gains are larger for shorter codes and increase as the number of perturbation branches grows.
Challenges improve speech extraction in multi-speaker conversations
Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement
Abstract: Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.