Sparse graph method improves audio-visual speech cleaning efficiency

G-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement

Computation and LanguageSound

Summary

Cleaning up speech recorded in noisy places is hard, especially when using both sound and lip movement to understand speech. The authors created a lightweight system called SG-Mamba that uses a smart graph to connect sounds and images in a way that saves computing power while keeping good accuracy. Their method also keeps some original sound details to avoid losing voice quality. Tests show SG-Mamba works well even in tricky multi-speaker noisy settings and runs efficiently on modest hardware.

What this means in practice

  • For mobile app developers: Build mobile apps that enhance speech quality by efficiently combining audio and visual cues for better voice clarity in noisy environments.
  • For hearing aid designers: Design hearing aids that improve speech understanding by selectively filtering noise using combined audio and visual signals with low computation needs.

Authors

Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen

Abstract

Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.