Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion
2026-08-17 • Sound
Sound
AI summaryⓘ
The authors developed DiffM2A, a new method that improves how we convert sounds captured by irregular and sparse microphone arrays into Ambisonic audio, which helps represent 3D sound scenes. Their approach uses a geometry-aware technique to better handle different microphone shapes without causing too much noise or losing detail. They also apply a diffusion model that learns from the microphone signals to produce more accurate spatial audio. Tests show that DiffM2A works better than existing methods in preserving sound quality and directionality, even for microphone setups it hasn't seen before.
AmbisonicsSpherical HarmonicsMicrophone ArrayDiffusion ModelSpatial AudioInverse FilteringBoundary ConditionsRotational EquivarianceSound IntensityBinaural Cues
Authors
Xiang Zhou, Zhengqiao Zhao, Zhengding Luo, Wen Zhang
Abstract
Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.