Papers for

speech system developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Machine unlearning shows limits for automatic speech recognition privacy

Evaluating Machine Unlearning in ASR

Abstract: Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whether existing MU algorithms and evaluation tools are suitable for ASR. We apply several MU techniques to an ASR model, evaluating privacy-utility trade-offs for single-subject unlearning, then assess the best algorithm under sequential and simultaneous unlearning. Results show that gradient ascent-based algorithms achieve strong utility-privacy trade-offs, whereas more complex approaches over-unlearn samples, making them easier to identify as unlearned. This suggests standard privacy evaluations based on simple Membership Inference attacks are insufficient to reliably assess unlearning success, motivating improved evaluation methods for MU in ASR. Finally, we show that both sequential and simultaneous unlearning yield worse privacy and utility than single-subject unlearning, underscoring the need for unlearning constructions better suited to these settings.

Mon 28 SeptComputation and Language
The gist
Some laws say people can ask to have their personal data forgotten, which means computer programs should 'unlearn' that data. This paper looks at how well current unlearning methods work for speech recognition systems, which turn spoken words into text. The authors found that simpler methods do a better job balancing how much the system forgets and how well it still works. They also discovered common tests don’t always reveal if data was fully forgotten, especially when unlearning multiple pieces at once. This means future tools need better ways to check if the unlearning really works in speech systems.
Open → 2609.34092v1

Audio language models get smarter at refusing harmful commands

AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks

Abstract: Large audio-language models (LALMs) expand language models to process and interpret audio, but also expose them to heterogeneous audio jailbreaks. We ask whether successful jailbreaks reflect failures to recognize harmful intent or failures occurring after such recognition. Layer-wise probing reveals the latter: risk-related information remains decodable from intermediate representations, yet the internal risk signal fails to translate into refusal in later-layer processing. We identify this discrepancy as the risk-to-refusal gap. Building on this finding, we propose AEGIS, a detect-then-intervene defense whose mid-layer risk gate selectively activates downstream safety adapters. Across six LALMs and three heterogeneous audio jailbreak benchmarks, AEGIS reduces the average unsafe rate from 17.9% to 0.4%, while causing only a marginal increase in over-refusal on benign inputs. These results establish selective internal intervention as an effective path toward more robust refusal in LALMs. The code is available at https://github.com/azzzzliao/aegis-audio-defense.

Thu 24 SeptSoundCryptography and Security
The gist
Large audio-language models can understand and process spoken content but sometimes fail to reject harmful or unsafe audio commands, a problem called jailbreaks. The authors found these models recognize risky inputs inside but fail to act on that knowledge to refuse unsafe requests. They created AEGIS, a method that detects risky audio in the middle layers of the model and then triggers safety measures to stop undesired answers. This approach greatly reduced unsafe responses across multiple models and test sets without blocking many harmless requests.
Open → 2609.29287v1

Speech data improves pinpointing topics in long transcripts

Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts

Abstract: Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.

Fri 18 SeptComputation and Language
The gist
Long audio transcripts can be hard to search because they contain a lot of extra information. This paper shows how using hidden speech features from the automatic speech recognition process can help find exactly where a topic is discussed in the transcript. The authors combine these speech features with text analysis to better locate relevant sentences without needing extra audio processing. Their method works best on structured talks like lectures or interviews but less well on casual conversations.
Open → 2609.21844v1

Speech llms adapted better to new voices with encoder adapters

Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs

Abstract: Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting them to domain-shifted speech, such as child or dialectal speech, remains challenging under limited target-domain data. Given the dominant role of the LLM in Speech-LLMs, with cross-entropy loss applied only at the LLM output, the speech encoder may receive insufficient adaptation to new acoustic conditions. In this paper, we propose Encoder Awakening via Adapters (EAVA), a simple yet effective domain-adaptive fine-tuning method for Speech-LLM-based ASR. First, lightweight adapters are inserted into each encoder layer and trained exclusively, enabling target-domain acoustic knowledge to be incorporated into the encoder while preserving its pre-trained knowledge. Second, the full model is jointly fine-tuned on the target domain with LoRA applied to the LLM. Experiments on three domain-shifted ASR datasets, covering child and dialectal speech, show that EAVA consistently outperforms vanilla fine-tuning and other baselines, achieving new state-of-the-art performance.

Wed 16 SeptComputation and Language
The gist
Speech recognition systems that use large language models struggle to understand voices different from those they were trained on, like kids or people with dialects. The authors propose adding small adapter modules inside the speech processing part of the system to better learn these new voice patterns without forgetting what was learned before. They also fine-tune the whole system together using a proven technique for the language model part. Their method works better than standard approaches when tested on challenging speech types.
Open → 2609.17981v1