Audio models tested for accuracy in stating numeric acoustic values
AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth
SoundComputation and LanguageMachine Learning
Summary
It can be hard to tell if a number spoken by an audio model about sounds is correct because people don’t always agree. The authors created AcoustiClaim, a way to check these numbers against precise instrument measurements. They tested different models on recorded speech sounds and found most had trouble giving accurate numbers consistently. Their system can also decide when to not give an uncertain number, which helps lower mistakes on mixed audio. This benchmark is a step toward understanding and improving how AI talks about sound features.
What this means in practice
- •For speech technology developers: Evaluate and improve models that estimate acoustic voice features by comparing their numeric outputs directly against instrument readings.
- •For audio signal processing engineers: Use AcoustiClaim to benchmark numeric accuracy of systems reporting measurements like pitch and shimmer on clean and mixed speech signals.
Authors
Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj
Abstract
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets' rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.