Equivariant deep learning improves sound event detection and localization efficiency

EquiSELD: Efficient training of equivariant sound event localization and detection networks

SoundArtificial Intelligence

Summary

Sound event detection and localization involves figuring out what sounds are happening and where they come from. The paper introduces a new method called EquiSELD that uses a mathematical symmetry property of sound signals to process them more efficiently and accurately. This approach splits the sound data into parts that stay the same under rotation and parts that change predictably, helping the model understand sounds regardless of their direction. Their method works better and trains faster than previous models using similar ideas, and also beats non-specialized models on simulated real-world sounds.

What this means in practice

  • For audio engineers: Design sound detection systems that accurately locate and identify sounds with less training time by using the EquiSELD model on spatial audio inputs.
  • For robotics developers: Equip robots with improved auditory perception to detect and localize environmental sounds under rotations and reflections using EquiSELD’s efficient equivariant processing.

Authors

Goksenin Yuksel, Marcel van Gerven, Kiki van der Heijden

Abstract

First-order Ambisonics (FOA) signals exhibit exact O(3) symmetry: The rotation or reflection of the FOA signal modifies the direction of arrival of the sound sources, while preserving the sound sources themselves. Prior attempts to utilize this spatial symmetry of FOA to improve the efficiency and robustness of sound event detection and localization (SELD) systems either learned only an approximation of the symmetry through rotation-based augmentation or relied on computationally expensive methods to integrate equivariance. Furthermore, prior work focused exclusively on SO(3) equivariance , leaving the potential of incorporating O(3) equivariance for SELD tasks unclear. To address these limitations, we developed EquiSELD. This equivariant attention network processes first-order Ambisonics as paired streams of O(3)-invariant scalars and equivariant intensity vectors, producing an invariant activity magnitude and an equivariant DOA with a Multi-ACCDOA readout. To compare the impact of O(3) versus SO(3)-equivariance, we designed a matched SO(3)-only variant. EquiSELD outperforms prior equivariant networks on both simulated scenes with measured RIRs and recordings of real-world sound scenes at a fraction of the training cost. EquiSELD additionally surpasses the performance of non-equivariant SELD networks of a similar size on the simulated real-world sound scenes and achieves competitive performance on the real-world sound scenes.