Papers for

audio signal processing engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Audio models tested for accuracy in stating numeric acoustic values

AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth

Abstract: Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets' rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.

Thu 24 SeptSoundComputation and LanguageMachine Learning
The gist
It can be hard to tell if a number spoken by an audio model about sounds is correct because people don’t always agree. The authors created AcoustiClaim, a way to check these numbers against precise instrument measurements. They tested different models on recorded speech sounds and found most had trouble giving accurate numbers consistently. Their system can also decide when to not give an uncertain number, which helps lower mistakes on mixed audio. This benchmark is a step toward understanding and improving how AI talks about sound features.
Open → 2609.30483v1

Speaker distance estimates improve with few real labeled examples

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Abstract: Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.

Thu 24 SeptSound
The gist
Measuring how far someone is from a microphone is tricky because models trained on simulated sounds don't work well with real recordings. The authors found that guessing the average distance is often better than using these models directly. However, by using just a small number of real recordings with known distances, they can adjust the model outputs to get better results without retraining. This process focuses on keeping the order of distances correct rather than perfect accuracy. Their work helps decide how many real examples are needed to improve these distance estimates effectively.
Open → 2609.29203v1