VoiceNet improves fine-grained understanding of voice emotions and styles
VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale
SoundArtificial Intelligence
Summary
Most existing systems can only recognize a few basic emotions from voices, often using acted speech that is less natural. The authors created VoiceNet, a new dataset that includes detailed human annotations for 40 emotions and over 50 voice style attributes in real-world speech recordings. They also developed VoiceCLAP models that learn to understand voice and text together, outperforming previous methods at identifying these fine-grained voice features. This work helps build tools that can better recognize subtle and complex voice expressions beyond simple emotion categories.
What this means in practice
- •For voice assistant developers: Filter and categorize large speech datasets into diverse voice emotions and styles to improve assistant responsiveness and personalization.
- •For call center quality teams: Automatically assess nuanced vocal attributes and emotions in customer calls to enhance training and customer experience monitoring.
Authors
Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby, Felix Friedrich, Maurice Kraus, Gijs Wijngaard, Kourosh Nadi, Huu Nguyen, Kristian Kersting, Sören Auer
Abstract
Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acted speech. This paper introduces VoiceNet, a human-annotated representation-level benchmark for voice performance understanding on permissively-licensed in-the-wild speech. VoiceNet has two subsets: VoiceNet-Emo applies a 40-emotion taxonomy with three expert ratings per item, and VoiceNet-Ext, a preliminary subset, scores 57 talking-style attributes including speaking rate, vocal tension, breathiness, and register. The paper also releases Emolia, an emotion-annotated version of the Emilia corpus, with a curated rebalanced subset enriched by dense MOSS-Audio Thinking annotations. Two voice-text contrastive baselines train on this data: a 110M-parameter VoiceCLAP-Small for fast large-scale data filtering and a 7B VoiceCLAP-Large for state-of-the-art performance. Both outperform existing CLAP baselines, which sit near chance on VoiceNet-Emo. On VoiceNet-Emo, VoiceCLAP-Large aligns more closely with the aggregate expert consensus than individual experts agree with one another: a comparison against the majority label rather than evidence of surpassing human emotion perception. All systems evaluated here are voice-text embedding models: VoiceNet scores representation-level attribute recognition and retrieval, not end-to-end spoken-dialogue behaviour. Clustering and filtering uncurated speech corpora into subsets that span diverse talking styles and emotions remains an open challenge; VoiceCLAP embeddings offer a promising tool for this task. VoiceNet, Emolia, and VoiceCLAP are publicly available for research use.