Language-Statistical Analysis of Neural Audio Codec Tokens Across Architectures, Corpora, and Noise Conditions
2026-08-31 • Computation and Language
Computation and LanguageSound
AI summaryⓘ
The authors studied how different types of neural audio codecs (NACs)—which turn speech into token sequences—behave statistically across various noisy conditions. They tested 13 different NAC designs on clean and noisy speech and measured patterns using language-analysis tools like Zipf's and Heaps' laws. They found that noise and codec type affect token statistics more than the speech dataset itself. Some degradation patterns previously seen in certain codecs were confirmed and extended to others. Their work helps guide how to analyze NAC tokens with language-like statistics depending on the codec architecture.
Neural audio codecResidual vector quantizationVector quantizationZipf's lawHeaps' lawUnigram entropyJensen-Shannon divergenceMel-cepstral distortionDEMAND noise
Authors
Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito, Hiroshi Saruwatari
Abstract
Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statistical laws. This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization (RVQ), single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions. Zipf and Heaps parameters, unigram entropy, codebook occupancy, and Jensen-Shannon divergence (JSD) are estimated from matched token samples with explicit fit-validity safeguards and family-conditional $n$-gram orders. Corpus identity explains little variance in any metric, whereas acoustic condition and quantizer meta-category dominate in a metric-dependent way, and unigram entropy is the metric most strongly associated with meta-category. Clean-to-noise JSD computed at a common unigram order is associated with mel-cepstral distortion most clearly under DEMAND noise. The collapse and explosion degradation signatures previously reported for RVQ codecs concentrate in RVQ cells under white and DEMAND noise, respectively; explosion also occurs in non-VQ codecs, and single-codebook VQ codecs shift in occupancy and distribution shape without either signature. These results provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.