Pretrained speech models enable flexible emotion editing without new training
SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows
SoundArtificial Intelligence
Summary
Editing the emotions in recorded speech usually requires lots of specialized training, which can be unstable. The authors explore whether large pre-trained text-to-speech models can be adjusted directly to change emotions without extra training. They find these models can indeed be changed to edit emotions but depend on the model type and editing steps to avoid unwanted effects. To make this practical, they created SEmoEdit, a method that edits speech emotions smoothly using the model's own internal flow, without retraining. They also built a benchmark dataset to test this approach and showed it outperforms other methods.
What this means in practice
- •For voice assistant developers: Integrate robust emotion control into speech synthesis without costly retraining using pretrained speech flows.$Commercial implications: This enables new emotion-aware voice products at scale by editing emotions in speech on demand using pretrained models without extra training.
- •For audio software engineers: Add dynamic, training-free speech emotion editing features into audio tools through direct manipulation of pretrained TTS models.
Authors
Tianxin Xie, Pengfei Zhang, Kai Jiang, Zelin Zhao, Li Liu
Abstract
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.