Frequency tuning improves text to speech sound quality and speed
Harmonizing Spectral Evolution in Conditional Flow Matching for TTS
Machine LearningSound
Summary
Generating speech from text can sometimes produce uneven or unnatural-sounding audio because low and high frequencies don't develop smoothly together. The authors found that by selectively boosting certain frequency ranges using a wavelet transform during the speech generation process, they can make the audio sound better and faster to produce. Their method works across different text-to-speech models, reducing the computer steps needed while improving audio quality without hurting speaker likeness or clarity.
What this means in practice
- •For speech synthesis engineers: Improve efficiency and audio quality of text-to-speech systems by tuning frequency evolution without retraining models.
- •For voice assistant developers: Speed up voice generation and enhance sound quality in AI assistants by applying frequency-selective boosting techniques.
Authors
Isha Pandey Varad Deshpande Abhijat Bharadwaj Ganesh Ramakrishnan
Abstract
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.