Speech language models improved for emotion recognition with linear classifier
Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
Computation and LanguageArtificial IntelligenceSound
Summary
Detecting emotions in speech is tricky because existing models often guess wrong or use labels not in the expected set. The authors changed how a speech language model reads emotions by adding a simple layer that picks the emotion directly from internal data without changing the main model. This approach reduces mistakes and works better, especially when using text from automatic speech recognition, revealing how subtle language biases connect with emotions. It also helps understand which words the model links to each emotion.
What this means in practice
- •For voice assistant developers: Improve emotion understanding in assistants by using a direct classification layer that reduces wrong emotion predictions from speech input.$Commercial implications: Enables more accurate emotion detection production features in voice-enabled devices and services increasing user satisfaction.
- •For call center analytics teams: Enhance emotion detection on transcribed customer speech for better insights without retraining core speech models.
Tested on one dataset.
Authors
Hasindri Watawana, Sergio Burdisso, Esaú Villatoro-Tello, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
Abstract
SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.