Speech enhancement model adapts to new acoustic environments without forgetting

Domain-Incremental Learning for Generative Speech Enhancement

Sound

Summary

Speech enhancement helps make voice recordings clearer by reducing noise, but models trained on one environment often struggle in new, different environments. The authors created a method that lets a single model learn to improve speech in many different noisy settings, one after another, without losing what it learned before. By using a special technique that updates only small parts of the model for each new environment, it remembers how to work well across all the places it has seen. This approach works better than just fine-tuning the whole model or hoping it handles new noise without extra training.

What this means in practice

  • For speech technology developers: Enable models to improve speech quality in new noisy environments while retaining performance on previous ones for better real-world deployment.
  • For call center engineers: Adapt voice enhancement systems incrementally to diverse call conditions without retraining from scratch or losing prior noise-removal capabilities.

Authors

Manjunath Mulimani, Annamaria Mesaros, Minje Kim, Jesper Rindom Jensen

Abstract

We propose a domain-incremental learning framework for generative speech enhancement (SE) that learns from a sequence of datasets or domains recorded under diverse acoustic conditions. Fine-tuning a pretrained model on continuously evolving domains leads to catastrophic forgetting of previously acquired knowledge, while zero-shot generalization often fails to adequately adapt to unseen domains. To address these challenges, we first develop a novel language model-based generative SE model that we then use as a pretrained backbone and incrementally adapt it to acoustically mismatched domains using lightweight domain-specific Low-Rank Adaptation. The proposed framework enables the model to acquire enhancement capabilities for new domains while preserving performance on previously learned domains. Evaluated on four heterogeneous speech datasets, our approach effectively adapts to new domains without forgetting previously learned domains.