Spatial audio models improve sound event detection and location accuracy
SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former
SoundArtificial Intelligence
Summary
It is hard for current computer models to tell what sounds are coming from where, especially when multiple sounds happen at once. The authors introduced a new model called SAIL that keeps the sounds and their locations separate inside the computer. This lets the model better recognize sounds, figure out their direction and distance, and understand how they relate in space. Their tests showed this method works better than older approaches that mixed sound and location information too early.
What this means in practice
- •For wearable device developers: Build devices that can recognize and locate multiple sound sources for enhanced user interaction and awareness.$Commercial implications: Enables smart wearables to provide spatial sound awareness and source identification, improving user experience and situational understanding.
- •For virtual reality developers: Improve immersive environments by accurately detecting and placing multiple sound sources in 3D space.
Authors
Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou, Woon-Seng Gan, Wenwu Wang
Abstract
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.