Papers for

assistive technology engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Audio visual system predicts speaker turns to isolate voices online

Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

Abstract: In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: https://jjjjiaozi.github.io/TS-VAP/.

Fri 25 SeptSound
The gist
Listening to one person speaking clearly during a conversation is hard when others talk at the same time. The authors created a new test setup using real two-person talks with unrelated background noise, rather than artificial mixtures. They designed a way to predict when a person will start or stop talking by using language understanding and facial cues. This prediction helps separate the right voice from all sounds quickly while the conversation happens. Their experiments show that guessing future talk times helps improve the clarity of the chosen speaker’s voice.
Open → 2609.30774v1

BranchShine-CR improves multilingual phonetic transcription accuracy efficiently

BranchShine-CR: Compact Multilingual IPA Transcription with Self-Conditioned CTC and Consistency Regularization

Abstract: We introduce BranchShine-CR, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet (IPA). It combines log-mel features, a rotary-position E-Branchformer encoder, intermediate self-conditioned connectionist temporal classification (CTC), and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances, it achieves 4.47% IPA character error rate, a 22.3% relative reduction from ZIPA-CTC-NS, with approximately one-twelfth as many parameters while being trained from scratch. BranchShine-CR also outperforms a similarly sized NeMo Conformer baseline across all 41 dataset language labels. Ablation studies indicate the individual components synergetically acting in model performance contribution. These findings support compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.

Thu 24 SeptMachine Learning
The gist
Transcribing many languages into a universal set of speech sounds called IPA is hard, especially on small devices. The authors introduce BranchShine-CR, a compact machine-learning model that understands many languages and transcribes their sounds more accurately than previous models with fewer parameters. It uses a combination of new techniques to stay accurate while being lightweight. This makes it easier to add detailed speech recognition into small devices even when computing resources are limited.
Open → 2609.29069v1

Audio visual speech recognition struggles outside broadcast settings

AVSRBench: A Multi-Condition AVSR Benchmark

Abstract: While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true generalization or just domain adaptation. To investigate this gap, we evaluate three AVSR architectures across six conditions: controlled broadcast speech, fixed-grammar utterances, hyper-articulated Lombard speech, read speech from professional lipspeakers and non-professional speakers, and spontaneous multi-party video conversations. We find that visual-only performance deteriorates rapidly beyond broadcast domains, and audio-video fusion mainly benefits Lombard speech environments. Visual understanding degrades sharply at 90° profile views, with multimodal systems relying largely on acoustic fallback. Additionally, speaker articulation proves more critical than minor camera shifts, and LLM-based architectures suffer from poor out-of-domain generalization. Our work highlights a significant generalization gap in current AVSR research. To address this, we also introduce RoomReader-AV as a new benchmark for AVSR and release a unified data preprocessing pipeline to make comprehensive multi-condition evaluation accessible.

Wed 9 SeptComputer Vision and Pattern RecognitionMultimedia
The gist
Speech recognition that uses both sound and lip movements works very well on TV broadcast speech, but the authors found it struggles when used in everyday situations like casual conversations or unusual speaking styles. They tested different systems in a variety of conditions and found that visual-only recognition fails quickly outside broadcast video, and even combined audio-visual systems mostly rely on sound when the speaker is not facing the camera. This shows current speech recognition technology does not generalize well to real-world video settings. The authors also offer a new dataset and tools to help researchers better test their systems in diverse conditions.
Open → 2609.10366v1