Audio source ID is less reliable after common sound file changes
Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution
SoundMachine Learning
Summary
Systems designed to identify which computer made an audio clip work very well in perfect conditions. The authors show that when audio files are compressed or changed in normal ways for storage or sharing, these systems lose a lot of accuracy. Different methods and audio conditions cause very different drops in performance, meaning clean test results don’t predict real-world performance. The study found that no single quality measure tells the whole story about how robust these systems will be after audio is altered. This means new tests considering these changes are important for real-world use.
audio provenance attributioncodectranscodingmacro-F1 scoreWavLMW2V2-BERTECAPA-TDNNproxy-anchorSI-SDRPESQ
Authors
Gang Shi
Abstract
Audio provenance attribution - which system produced a synthetic utterance - is reported at near-ceiling accuracy on clean benchmarks, yet audio reaching an analyst has usually been transcoded. We report a prospectively registered measurement of closed-set attribution after single-stage codec transport, with the analysis region fixed from fidelity metadata before any attribution model was trained. On two corpora, in-support losses reach 53.5 [43.5, 63.6] and 70.3 [63.0, 77.5] Macro-F1 points for WavLM-Base+, and 61.0 [56.8, 65.1] and 49.8 [41.6, 57.9] for W2V2-BERT 2.0, under simultaneous component-level bands. Degradation is strongly condition- and representation-dependent: within one in-support grid WavLM losses run from -0.4 to +53.5 points, and the two encoders differ beyond a prespecified +/-5-point margin at six of twelve conditions. A clean-qualified ECAPA-TDNN and a Proxy-Anchor head degrade comparably, so the effect is not confined to one representation family or a weak linear head. The registered matched-fidelity comparison was not estimable on this grid, and waveform and perceptual measures order the conditions differently: MP3 at 8 kbit/s ranks mid-grid on SI-SDR but last on PESQ-WB while causing the largest loss. For the tested tasks, corpora, representations and codec grid, a clean accuracy figure does not by itself characterise deployment robustness.