Papers for

telecommunication engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Target speech extraction improves with progressive adaptation to real speech

RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing

Abstract: Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.

Mon 28 SeptSound
The gist
Extracting a person's speech in noisy real-world conversations is hard because training mostly uses fake data that sounds different from real sounds. The authors adapted a method called RemixIT to teach computers progressively using both made-up and real conversation data without perfect labels. Their two-step process first mixes real and synthetic speech information, then refines learning using reliable guesses about real speech signals. This approach made it easier for the system to identify target speech accurately in real conversations, beating older methods in several tests.
Open → 2609.35118v1

Polar code decoding improved with offline variance perturbation design

Parallel Successive Cancellation Perturbation-Enhanced Decoding of Polar Codes via Offline Variance Design

Abstract: Successive cancellation perturbation-enhanced (SCP) decoding improves the performance of finite-length polar codes by performing multiple SC decoding attempts with receiver-side perturbations. However, many existing perturbation schemes generate or update subsequent perturbations according to the outcomes of previous decoding attempts, resulting in additional decoding latency. In this paper, we propose an offline variance design (OVD) method for parallel SCP (PSCP) decoding of short- and medium-length polar codes. First, we formulate the exact recovery objective conditioned on ordinary SC failure and classify failed frames by the position of the first genie-aided intrinsic error and the number of subsequent intrinsic errors. We also derive a consistent Gaussian representation of the perturbed channel that preserves min-sum SC hard decisions. Second, we construct a class-based approximation of the recovery objective using Gaussian approximation and backward recursions, accounting for both error correction and new errors introduced by perturbations. We prove that both the exact and analytical objectives are nondecreasing and exhibit diminishing marginal gains as independent branches are added. Third, we develop a greedy algorithm to select variances from a finite candidate set for a given code, signal-to-noise ratio (SNR), and number of perturbation branches. All variances are determined offline, allowing the original SC branch and all perturbation branches to start simultaneously. Simulations for rate-$1/2$ polar codes of lengths $64$, $128$, $256$, and $512$ show that OVD-PSCP achieves lower block error rates (BLERs) than conventional SCP with the same number of perturbation branches. The gains are larger for shorter codes and increase as the number of perturbation branches grows.

Mon 28 SeptInformation Theory
The gist
Polar codes help protect data sent over noisy communication channels. The authors show a new way to improve decoding these codes by preparing all the trial corrections ahead of time instead of waiting for previous attempts. This approach lets the decoder try several possibilities in parallel, making it faster and more effective, especially for shorter messages. Their method uses math to pick the best variations before decoding starts, which reduces errors when decoding data.
Open → 2609.34059v1

Challenges improve speech extraction in multi-speaker conversations

Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

Abstract: Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.

Tue 22 SeptSoundComputation and Language
The gist
In noisy conversations with many people, pulling out one person’s voice clearly is hard, especially when speakers talk less and their reference samples differ from what they say in the conversation. The authors found that current methods trained on balanced, simulated data don’t work as well in real-life talks. They created a new training technique to handle long silences better, improving the clarity and quality of the extracted speech. They also studied how mismatched reference samples affect performance.
Open → 2609.25948v1