EmoRES improves emotional speech control in text to speech systems

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

SoundComputation and Language

Summary

Sometimes, computer systems that turn written text into spoken words struggle to express emotions correctly. The authors found that an emotion signal inside these systems can be broken into two parts: one that just changes speech away from neutrality, and another that steers it toward a specific emotion. They created a method called EmoRES that adjusts these parts separately without needing extra training. This approach helped machines produce speech that better matches emotions and sounds more natural, according to tests and listener feedback.

What this means in practice

  • For voice assistant developers: Generate emotionally expressive speech outputs without additional model training to enhance user interaction.
  • For game sound designers: Create character voices with controllable emotions quickly using existing speech models and vector adjustments.

Authors

Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, Neha Chachra

Abstract

Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.