Speech content features compared on audio generation and identity control

A Comprehensive Study of Content Representations for Speech Synthesis

Machine LearningSound

Summary

It can be hard to tell how different ways of representing spoken language actually affect the sounds that a computer makes when it tries to copy or change someone's voice. This study trained a computer model to recreate speech using different types of speech representations, then checked how well each kept the words, the speaker’s unique voice, and the rhythm of speech. The authors found that some representations do a great job of copying the original sound, while others better separate who is speaking from what is being said. They showed that whether or not the speaker’s identity can be separated depends not just on the method used, but also on how much information the representation can hold.

What this means in practice

  • For speech technology developers: Build voice conversion systems that better separate who is speaking from what is said by choosing appropriate content representations.
  • For multimodal ai engineers: Design speech-to-speech translation tools that maintain speaker identity and speech content by applying insights on representation capacity and training objectives.

Authors

Diego Torres, Axel Roebel, Nicolas Obin

Abstract

Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.