SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching
2026-08-03 • Sound
Sound
AI summaryⓘ
The authors address the problem of separating a target speaker's voice from others when a listener's attention switches between speakers during a trial, which is hard because brain signals are noisy and delayed. They propose SAGE, a new method that creates two possible speech versions and smoothly combines them based on EEG signals, reducing glitches when attention shifts. Their approach also adjusts for timing delays and uncertainties in brain data, improving accuracy and reducing response time. Overall, their method better tracks who the listener is focusing on and extracts that speaker's voice more reliably.
EEGauditory attentiontarget speaker extractionspeech separationsignal latencysoft gatingSI-SDRSTOIneural decodingattention switching
Authors
Xuefei Wang, Ximin Chen, Yuting Ding, Chunlin Li, Fei Chen
Abstract
EEG-guided target speaker extraction is challenging under in-trial auditory attention switching, where neural noise and intrinsic latency can delay or destabilize attention tracking. Conventional methods struggle with dynamic switches and often cause discontinuities at switching points. Therefore, we propose SAGE, a switch-aware EEG-guided soft gating framework that treats in-trial switching as dynamic selection. SAGE generates two candidate speech streams with a robust separator and uses an EEG-guided switch-aware gating module to produce smooth fusion weights and suppress transition artifacts. We further integrate latency-compensated alignment and an uncertainty-driven conservative strategy to handle latency discrepancies and fluctuating EEG reliability. SAGE outperforms baselines, achieving 8.67 dB SI-SDR and 88.24% STOI while reducing average switching latency to 2.04 s. By coupling neural decoding with speech separation, it enables robust target extraction in dynamic scenarios.