Speech models struggle with technical talk in science fields
$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants
Computation and Language
Summary
Voice assistants are good at general conversations but have trouble in scientific areas where people use complex terms and symbols. The authors created S3-Bench, a way to test how well these voice assistants handle science topics across 10 different fields. They look at speech recognition, understanding, reasoning, and speaking responses, finding that current models often fail to adapt well or give complete and accurate answers during longer talks. This shows that more work is needed to make voice assistants better for specialized scientific use.
What this means in practice
- •For voice assistant developers: Evaluate and improve voice assistants' performance on specialized scientific conversations using S3-Bench’s framework.
- •For customer experience teams: Assess and enhance multi-turn speech interactions in technical domains to reduce user frustration and improve accuracy.
Authors
Heyang Liu, Jiayi Huang, Wenyang Xiao, Ziyang Cheng, Lixin Zhang, Zhen Liu, Miao He, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang
Abstract
The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplored, particularly in scientific areas. Scientific interactions introduce formidable challenges, involving rare technical terminology, spoken norms of abbreviations, and the natural verbalization of symbolic special expressions. In this paper, we introduce S$^3$-Bench, a systematic evaluation framework covering 10 major disciplines, consisting of a Knowledge set for speech question-answering and a Dialogue set for multi-turn progressive interactions with simulated user agents. By decomposing a complete atomic turn into stages of speech recognition, perception, knowledge utilization with reasoning, and response pronunciation, we systematically characterize the common challenges and performance tradeoffs of existing approaches. Furthermore, experiments on multi-turn interactions reveal persistent limitations in user adaptation and the generation of accurate, comprehensive, and efficient responses.