Text to audio generation improved by joint learning of sounds and models
UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation
Sound
Summary
Creating audio from text usually involves first turning sounds into a code and then generating new sounds from that code. The authors found that making the code and the sound generation model learn together leads to better results. Their new method, called Unite-Audio, improves how natural the generated audio sounds based on text descriptions. They also added a way to fine-tune the model after training to produce even better audio.
What this means in practice
- •For audio software developers: Create more natural-sounding audio from text by training representation and generation together for better audio quality.
- •For game sound designers: Generate diverse and realistic sound effects from textual descriptions more efficiently with a compact audio generation model.
Authors
Runwu Shi, Kai Li, Yujin Wang, Dong Yang, Jiahui Li, Jiang Wang, Benjamin Yen, Ashizawa Takeshi, Chunxiang Jin, Kazuhiro Nakadai
Abstract
Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. However, reconstruction-oriented representations may be suboptimal for generation, motivating joint representation and generative learning. To this end, we introduce \textbf{Unite-Audio}, to our knowledge, is the \textbf{first} to jointly learn continuous audio representations and latent flow matching for TTA. By coupling reconstruction with self-supervised generative prediction, Unite-Audio allows the generative objective to directly shape the latent space rather than treating it as a fixed intermediate representation. We further employ Flow-GRPO post-training to improve text-conditioned generation. Experiments show competitive TTA performance with a compact latent flow model, while ablation studies confirm the benefit of jointly learning the audio representation and generative model. Audio samples are available at https://runwushi.github.io/Unite-Audio.