Speech models struggle to interpret non-speech vocal emotions accurately

NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models

Computation and LanguageHuman-Computer Interaction

Summary

People often use non-speech sounds like laughs or sighs in conversations to express feelings. This paper by the authors checks if speech-to-speech computer models can notice these sounds and respond correctly. They created a special test with pairs of talks that are the same except for the non-speech sounds at the end. Their tests show models can usually tell when these sounds happen but have trouble understanding their emotional meaning and reacting differently. The authors also shared their test data and tools for others to use.

What this means in practice

Authors

Ziwei Chen

Abstract

We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final turn. Our pilot contains 22 human-verified pairs (44 audio conditions) and evaluates five models on NSV perception, emotion understanding, and response adaptation. Results show that models generally perform better at detecting NSVs than at interpreting their fine-grained emotional meaning or producing appropriately differentiated responses. The data construction pipeline, dataset, and evaluation pipeline are publicly available at https://github.com/ChenzwNina/nsv-construction.