Audio encoders balance size and speed for better speech tasks

Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders

Sound

Summary

Audio encoders shrink sound data by changing two things: how many features they keep (width) and how often they sample in time (frame rate). The authors trained many versions with different widths and rates to see how these choices affect tasks like speech recognition and answering spoken questions. They found bigger isn’t always better; moderate sizes at higher speeds worked best for understanding speech, even if bigger sizes gave better sound reconstruction. This shows that just making the audio sound closer to the original doesn’t guarantee better performance in real tasks.

What this means in practice

  • For voice assistant developers: Optimize audio encoder settings to improve speech recognition accuracy in voice-controlled devices by balancing feature size and sampling speed.
  • For speech data engineers: Design audio feature extraction pipelines that maintain good task performance despite compression by adjusting dimensionality and frame rates together.

Authors

Kyudan Jung, Sehyun Lee, Son-ha Jo, Jaegul Choo, Sanghyuk Shoi

Abstract

Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best observed widths shifting toward larger values under stronger temporal compression. Frozen-model PCA interventions reveal distinct reconstruction and recognition sensitivities: removing the trailing half of the components substantially degrades ASR in high-rate 512-dimensional encoders with comparatively small reconstruction penalties, whereas 1024-dimensional encoders largely preserve both. Yet the projected 1024-dimensional model underperforms unmodified narrower models on ASR at 12.5Hz. These findings identify a width--rate interaction in downstream utility and suggest that how representations are organized during training matters beyond reconstruction fidelity and compressibility.