Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

2026-08-10Artificial Intelligence

Artificial Intelligence
AI summary

The authors created a new test called ST-OmniQA to see how well computers understand moving sounds in videos, like identifying what is making the sound, where it is, and how it moves. They used special 360-degree videos paired with detailed 3D audio recordings to make questions that check these abilities. Then, they made a model named ST-Omni-R1 that combines sound and visual information and learns step-by-step to answer these questions better than existing models. Their model also works well on other similar sound-location tasks, showing it learns useful ways to understand sound and movement.

spatio-temporal audio-visualAmbisonicspanoramic videosound source localizationaudio-visual reasoningcurriculum learningreinforcement learningsemantic accuracydirection of arrivalmotion trajectories
Authors
Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang, Yifei Zheng, Minnan Luo
Abstract
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.