Speech llms adapted better to new voices with encoder adapters
Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs
Computation and Language
Summary
Speech recognition systems that use large language models struggle to understand voices different from those they were trained on, like kids or people with dialects. The authors propose adding small adapter modules inside the speech processing part of the system to better learn these new voice patterns without forgetting what was learned before. They also fine-tune the whole system together using a proven technique for the language model part. Their method works better than standard approaches when tested on challenging speech types.
What this means in practice
- •For speech system developers: Create better speech recognition systems that quickly adapt to new speaker types like children or dialect speakers with limited training data.
- •For customer support teams: Improve transcription accuracy for varied customer accents and dialects by deploying speech models fine-tuned using encoder adapters.
Authors
Mohan Shi, Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang, Eray Eren, Abeer Alwan
Abstract
Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting them to domain-shifted speech, such as child or dialectal speech, remains challenging under limited target-domain data. Given the dominant role of the LLM in Speech-LLMs, with cross-entropy loss applied only at the LLM output, the speech encoder may receive insufficient adaptation to new acoustic conditions. In this paper, we propose Encoder Awakening via Adapters (EAVA), a simple yet effective domain-adaptive fine-tuning method for Speech-LLM-based ASR. First, lightweight adapters are inserted into each encoder layer and trained exclusively, enabling target-domain acoustic knowledge to be incorporated into the encoder while preserving its pre-trained knowledge. Second, the full model is jointly fine-tuned on the target domain with LoRA applied to the LLM. Experiments on three domain-shifted ASR datasets, covering child and dialectal speech, show that EAVA consistently outperforms vanilla fine-tuning and other baselines, achieving new state-of-the-art performance.