AcoustiTrace: When Plausible Sound Violates Physics
2026-08-03 • Multimedia
MultimediaSound
AI summaryⓘ
The authors created AcoustiTrace, a new way to check if audio-video generators produce sounds that follow real-world acoustic rules, like how sound travels or how it is heard. They developed a large dataset and tests focusing on eight specific sound properties to better understand where current models go wrong. Their findings show that even top models make mistakes with basic acoustic details, despite sounding realistic. AcoustiTrace can help improve these models by pointing out exactly what needs fixing to make generated sounds more physically accurate.
audio-video generationacoustic realismsound propagationevaluation benchmarktext-to-audio-videoimage-to-audio-videoRGB-D observationsacoustic processessound generationmodel refinement
Authors
Shiyang Li, Yuewen Cao, Yihao Liu, Yuandong Pu, Baochang Zhang, Xiaofei Li, Changqing Zou
Abstract
Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.