DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
2026-08-10 • Sound
SoundArtificial Intelligence
AI summaryⓘ
The authors created DAVE, a system to make speech clearer when both sound and video are used, especially in tough real-life situations where video quality can be bad and training data is limited. They built a huge dataset called DAVE-Corpus to help train the system better. Their approach improves different aspects of speech quality, like how clear and natural it sounds, by using a step-by-step training method. They also designed a smart process to only enhance parts that need fixing, avoiding any harm to good-quality audio. Tests show that DAVE works well even when videos are unclear or sounds are mixed.
audio-visual speech enhancementspeech separationtraining corpusmulti-objective optimizationGAN-based denoisingloudness normalizationvisual degradationReal-World Audio-Visual Speech Enhancement Challenge
Authors
Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
Abstract
Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions. Existing approaches usually fuse visual features directly into the separation network, making them vulnerable to degraded visual signals. In this paper, we present DAVE, a decoupled audio-visual enhancement framework for real-world speech separation. Firstly, to address the data scarcity issue, we construct DAVE-Corpus, a large-scale training corpus with 219,411 mixtures generated from public meeting corpora through combinatorial acoustic augmentation. Then, we introduce a progressive multi-objective optimization strategy to jointly improve speech separation, intelligibility, speaker identity preservation, and perceptual quality. We further develop a certified selective enhancement chain that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics. Experimental results on the Real-World Audio-Visual Speech Enhancement Challenge demonstrate the robustness of DAVE under both real-world mixed scenarios and visual degradation conditions.