Objects as Audio-Visual Modal Sound Fields

2026-08-05Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors created a new way to guess what objects sound like when you hit them, using just a few real sounds and pictures from different angles. They combined 3D models of the object’s shape with key sound features to make a simple but effective sound model. This method works better than older ways that use lots of data or complicated simulations. Their tool can also help find where something was touched and let people change the sounds objects make.

3D reconstructionacoustic cuesimpact soundGaussian Splattingmodal sound fieldfew-shot learningcontact localizationsound renderingvisual feature integration
Authors
Zisen Shao, Zihao Wei, Derong Jin, Ruohan Gao
Abstract
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.