SyncRA improves timing connections between sounds and images in videos
SyncRA: Learning Temporal Correspondence in Omni-Modal Models
Computer Vision and Pattern RecognitionArtificial IntelligenceMultimedia
Summary
Many AI models that look at videos and listen to sounds have trouble matching the right sound to the right image at the same time. This can cause them to misunderstand what they see and hear together. The authors found that these mistakes happen because the models do not reliably link sound and visuals by their timing. They created SyncRA, a method that teaches models to better connect sounds and images happening together without needing extra labels or changing how the models work when watching videos. This method improved several popular AI models on tests where they had to answer questions about videos with both sound and pictures.
What this means in practice
- •For video analysis teams: Improve video content understanding by accurately linking sounds to their matching visual moments within videos.
- •For multimedia software developers: Enhance audio-visual question answering features by integrating timing-aware training that strengthens synchrony detection in existing models.$Commercial implications: Enables building commercial interactive video assistants that reliably understand combined audio and visual content, improving user experience.
Authors
Zelong Xu, Yan Li, Wenhe Hu, Xiyang Hu
Abstract
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.