Spoken sarcasm detection relies more on words than voice tone

CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection

SoundComputation and Language

Summary

Detecting sarcasm in spoken language can depend on the words people use, the way they say them (tone, speed), or both together. The authors created a way to test which of these clues speech recognition systems actually use to spot sarcasm. They found that the words themselves usually play a bigger role than the tone or prosody when identifying sarcasm in speech. This helps understand how sarcasm detection systems work and shows that changing tone alone doesn’t always change the system’s sarcasm judgment.

What this means in practice

Authors

Qiyang Sun, Xudong Li, Yupei Li, Jiabin Xue, Yuhang Dai, Jiaming Li, Bjorn W. Schuller

Abstract

Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under original, lexical-preserving, prosody-preserving, and approximately neutralised conditions. We evaluate handcrafted acoustic-feature systems, self-supervised learning (SSL) probes, and large audio language models (LALMs) on CMMA and MUStARD. For target-only Qwen3-Omni, lexical-preserving speech retains a 0.135--0.148 AUROC advantage over prosody-preserving speech after duration balancing, with cluster-bootstrap intervals above zero; alternative lexical resynthesis preserves this advantage. Acoustic interventions shift scores without consistently improving discrimination or changing binary predictions under the evaluated conditions. Context and interaction estimates vary across corpora. These findings distinguish acoustic sensitivity from sarcasm discrimination while exposing duration, identity, and transformation effects.