Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation

2026-08-10Sound

Sound
AI summary

The authors focus on improving how long speech recordings are separated into individual speakers. They propose a method that doesn't require extra training and aligns speaker segments by comparing speaker characteristics (embeddings) with a reference set. This method updates its references dynamically to keep the most accurate speaker information and works as an add-on to existing speech separation tools. Their approach works better than previous methods, especially when speech segments are far apart or when the number of speakers isn't known accurately.

speech separationspeaker embeddingclusteringcosine similaritypermutation alignmentpost-processingreference poollong speechunknown speaker countdynamic clustering
Authors
Yuzhu Wang, Archontis Politis, Konstantinos Drossos, Tuomas Virtanen
Abstract
Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its challenge lies in predicting cross-segment permutations. This paper proposes a training-free dynamic clustering approach for cross-segment permutation alignment using speaker embedding reference pools. The method predicts the permutation using the cosine similarity between current segment embeddings and the reference pools. The approach updates reference pools by retaining the most representative speaker embeddings based on their overall cosine similarity with existing references. As a plug-and-play post-processing module compatible with existing separation models, the proposed method demonstrates superior performance compared to existing methods on dense and sparse long speech scenarios, particularly in challenging sparse scenarios with extended utterance gaps, and further shows robustness to speaker count estimation errors in unknown speaker count scenarios.