FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting

2026-07-20Multimedia

Multimedia
AI summary

The authors address the challenge of creating realistic impact sounds based on what you see, focusing especially on how the inside filling of an object (like water or rice) changes the sound. They made a new dataset called FillImpact with over 5,000 sound recordings from different real objects, varying how much they're filled and what hits them. Their analysis shows these sounds follow real physical rules about sound resonance and damping. Using this data, they developed a new method called FillGauss that combines 3D shape information, the point of impact, and details about the filling to generate accurate impact sounds. Their work improves how well AI can create sounds that match real-world physics in a cross-modal way.

impact sound synthesis3D Gaussian Splattingacoustic resonancedampingmulti-modal AIlatent diffusioncross-modal generationphysical modelinginternal fillingsound generation dataset
Authors
Chen Yang, Ganye Wen, Bin Huang, Jiayi Lyu, Zehai Niu, Linlin Shen, Jinbao Wang
Abstract
Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling states, a critical physical factor that drastically modulates acoustic resonance and damping. To address this issue, we have defined a new task called Fine-Grained Filling-Aware Impact Sound Generation. As a foundational step, we first introduce the fine-grained fill-aware dataset (FillImpact), a pioneering multi-modal collection comprising over 5,000 rigorous acoustic recordings from 88 diverse real-world objects. It captures impact interactions with varying internal contents (i.e., water, rice), a continuous range of fill levels, and distinct striker materials. Furthermore, comprehensive acoustic analysis confirms that the collected data closely aligns with established physical laws governing acoustic resonance and damping, indicating its suitability for physically grounded modeling. Building on this dataset, we propose a novel generative framework (FillGauss) that integrates 3D Gaussian Splatting (3DGS) with internal state conditioning for sound generation. By fusing 3DGS geometric features, precise 3D spatial strike coordinates, and fine-grained textual physical conditions within a latent diffusion architecture, FillGauss enables position-aware, striker-aware, and filling-aware audio generation. Extensive experiments demonstrate that our approach could generate high-fidelity impact sounds that adhere to underlying physical principles, establishing a new state-of-the-art for physically grounded cross-modal audio generation.