Speaker distance estimates improve with few real labeled examples

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Sound

Summary

Measuring how far someone is from a microphone is tricky because models trained on simulated sounds don't work well with real recordings. The authors found that guessing the average distance is often better than using these models directly. However, by using just a small number of real recordings with known distances, they can adjust the model outputs to get better results without retraining. This process focuses on keeping the order of distances correct rather than perfect accuracy. Their work helps decide how many real examples are needed to improve these distance estimates effectively.

What this means in practice

  • For smart home device makers: Improve voice assistant responsiveness by adapting speaker distance estimates with minimal real-world labeled data.$Commercial implications: Enables more accurate distance sensing in commercial smart speakers by calibrating synthetic models to actual user environments.
  • For audio signal processing engineers: Calibrate existing synthetic-trained models quickly using few real samples to enhance distance estimation in room acoustics analysis.

Authors

Michael Neri, Archontis Politis, Tuomas Virtanen

Abstract

Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.