Large audio models use layered clues to tell when sounds happen
On Temporal Binding in Large Audio Language Models
SoundMachine Learning
Summary
Figuring out when sounds happen in audio recordings is tricky for AI models. The authors studied three large audio language models to see how they guess the timing of sound events. They found these models store rough timing information in parts of the model that combine sound and language data. Moving these timing clues slightly changes the model’s sense of when events occur, but it doesn’t affect the exact timing predictions. This means rough timing and precise timing use different parts of the model’s reasoning.
What this means in practice
- •For audio application developers: Improve audio apps by targeting model layers that track rough sound timing to enhance event order understanding.
- •For speech recognition engineers: Separate model improvements for coarse event timing and precise event timestamping to boost speech-to-text accuracy.
Authors
Paul Primus, Gerhard Widmer
Abstract
Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model components that may need improvement. Using mechanistic interpretability, we investigate how temporal information is represented and bound to sound events in three open-source LALMs. We find that across all three, event-specific location becomes concentrated in event name representations at intermediate modality integration layers. These representations encode coarse event position along a low-dimensional, curved relative time trajectory. Steering event name representations along this trajectory systematically shifts before/after beliefs, providing evidence that these representations contribute to coarse temporal reasoning. In contrast, the same interventions do not reliably shift predicted onset timestamps, suggesting that coarse temporal reasoning and precise metric event localization rely on distinct mechanisms.