Multimodal system improves emergency vehicle recognition with audio and video
Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation
MultimediaSound
Summary
Emergency vehicles like ambulances and fire engines are hard to spot by either sound or sight alone, especially in noisy or dark conditions. The authors created a system called AVNet that combines both audio and video to better recognize these vehicles. It cleverly pairs sounds and images from the same moments and can still work if one of the inputs (sound or video) is missing. Their tests show that combining both types of information helps detect emergency vehicles much more accurately than using either sound or video alone.
What this means in practice
- •For autonomous vehicle manufacturers: Improve safety systems by reliably detecting emergency vehicles using combined audio and visual inputs, even in challenging conditions like noise or darkness.$Commercial implications: Allows manufacturers to sell more robust emergency vehicle detection features that improve autonomous driving safety.
- •For urban traffic management teams: Deploy better surveillance to detect emergency vehicles accurately across noisy and visually obstructed city environments using combined audio-visual analysis.
Tested on one dataset.
Authors
Vijay John, Amar Dabaja
Abstract
Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems fail in noisy urban environments, and camera-based systems fail at night or under occlusion. This report presents AVNet, a multimodal audio-visual transformer that classifies emergency vehicles (ambulance, fire engine, police car) and road background using both audio and video, while gracefully handling the absence of either modality at inference time. AVNet introduces three key contributions: (1) a temporally aligned cross-modal fusion module that performs second-level cross-attention between audio spectrogram tokens and video frame tokens, exploiting their exact temporal correspondence without any learned alignment mechanism; (2) learned null embeddings that substitute for missing modality tokens, enabling a single unified model to operate in audio-only, video-only, or joint audio-visual mode without retraining; and (3) a knowledge distillation training strategy in which specialist unimodal teacher models transfer inter-class dark knowledge into the multimodal student fusion branch via soft probability targets. Evaluated on 281 clips from the Google AudioSet dataset, AVNet achieves 66.6% overall accuracy in audio-visual mode, outperforming the audio-only branch by +10.4% and the video-only branch by +15.0%. The largest per-class gain is observed for the hardest class, Ambulance, where fusion achieves +29.5% over either unimodal branch alone, demonstrating that the two modalities provide complementary information that the aligned cross attention mechanism successfully exploits.