Speech replay attack detection adapts to changing environments over time
Domain-Incremental Learning for Multi-Channel Replay Speech Detection
Summary
Replay attacks trick voice-controlled systems by playing back recorded speech, and detecting these attacks is harder because different rooms and places change how the sound is heard. The authors looked at how to make a detector learn to recognize these attacks in many different acoustic settings one after another without needing to keep all old recordings. They found that simply retraining forgets old environments quite badly, but using special methods that keep separate sound filters for each environment helps the detector remember better and perform more accurately. Their work tested this idea across many orders of environments, showing that the last place learned has a big impact on how well the detector works.
What this means in practice
- •For speech security engineers: Update replay speech detectors to adapt over time to different acoustic environments without storing all past data, improving long-term reliability.
- •For audio device manufacturers: Design multi-microphone devices with environment-specific acoustic processing to enhance replay attack detection performance across diverse settings.