Papers for

assistive technology teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Personalized Korean lipreading system cuts visual speech recognition errors

Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS

Abstract: We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer initialized from English-trained weights attains 9.95 - 12.19% character error rate (CER) under the corpus protocol against the published 26.64, and 19.00 - 21.52 on unseen wording. Per speaker, CER spans 1.0 to 52.2%, with seen wording lowering CER by 7.0 - 9.0 points and professional delivery and spontaneous speech raising it by 8.5 - 10.5 and 12.7 points. A low-rank adapter with 4.6% of the parameters, trained on 4 to 29 minutes of the user's frontal video, lowers the CER of twelve high-error speakers by 2.13 to 3.58 points, transfers to every camera without loss, and keeps 85% of the full fine-tuning gain at 12% of its cost to other speakers. Cameras above the mouth plane add about six CER points as a constant offset that training on all views keeps small.

Thu 24 SeptComputation and LanguageComputer Vision and Pattern Recognition
The gist
Visual speech recognition is like reading lips from videos, but it can make many mistakes. This paper shows a method to tailor a Korean lipreading system to individuals, greatly reducing errors even with only a few minutes of user video. The system also works well from different camera angles and copes with different speaking styles and unfamiliar words better than before. This helps close the gap between a general system’s average accuracy and how well it works for specific people.
Open → 2609.28988v1

Speech recognition learns new words itself from unlabeled test audio

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

Abstract: New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.

Thu 24 SeptComputation and Language
The gist
Speech recognition systems usually struggle with new words they have never heard before. This paper shows how such systems can learn new words by listening to unlabeled speech data during testing, much like humans learn new words from context. The method uses a fixed speech-to-text model to generate spelling guesses and a language model to confirm the context of possible new words. Together, an adaptation module improves recognition of these new words without needing manual transcripts. Testing showed it reduces errors significantly on both regular and impaired speech datasets.
Open → 2609.28877v1