Spoken sarcasm detection relies more on words than voice tone
CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
SoundComputation and Language
Summary
Detecting sarcasm in spoken language can depend on the words people use, the way they say them (tone, speed), or both together. The authors created a way to test which of these clues speech recognition systems actually use to spot sarcasm. They found that the words themselves usually play a bigger role than the tone or prosody when identifying sarcasm in speech. This helps understand how sarcasm detection systems work and shows that changing tone alone doesn’t always change the system’s sarcasm judgment.
What this means in practice
- •For speech technology developers: Design and evaluate sarcasm detection systems by distinguishing reliance on word choice versus speech tone cues.
- •For call center quality teams: Improve assessment tools by understanding if sarcasm recognition depends on voice prosody or word usage in customer interactions.
Authors
Qiyang Sun, Xudong Li, Yupei Li, Jiabin Xue, Yuhang Dai, Jiaming Li, Bjorn W. Schuller
Abstract
Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under original, lexical-preserving, prosody-preserving, and approximately neutralised conditions. We evaluate handcrafted acoustic-feature systems, self-supervised learning (SSL) probes, and large audio language models (LALMs) on CMMA and MUStARD. For target-only Qwen3-Omni, lexical-preserving speech retains a 0.135--0.148 AUROC advantage over prosody-preserving speech after duration balancing, with cluster-bootstrap intervals above zero; alternative lexical resynthesis preserves this advantage. Acoustic interventions shift scores without consistently improving discrimination or changing binary predictions under the evaluated conditions. Context and interaction estimates vary across corpora. These findings distinguish acoustic sensitivity from sarcasm discrimination while exposing duration, identity, and transformation effects.