Audio visual system predicts speaker turns to isolate voices online

Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

Sound

Summary

Listening to one person speaking clearly during a conversation is hard when others talk at the same time. The authors created a new test setup using real two-person talks with unrelated background noise, rather than artificial mixtures. They designed a way to predict when a person will start or stop talking by using language understanding and facial cues. This prediction helps separate the right voice from all sounds quickly while the conversation happens. Their experiments show that guessing future talk times helps improve the clarity of the chosen speaker’s voice.

What this means in practice

  • For conferencing system developers: Improve real-time isolation of a selected speaker’s voice during live video calls with competing speech and noise.$Commercial implications: Enables enhanced speaker clarity in commercial video conferencing products by predicting speaker turns for better audio separation.
  • For assistive technology engineers: Create hearing aids or augmented listening devices that better track who is speaking in complex, multi-person conversations.

Authors

Shuhan Zhang, Wenxuan Wu, Haizhou Li

Abstract

In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: https://jjjjiaozi.github.io/TS-VAP/.