Papers for

call center quality teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Speechcritic improves speech quality judgments using limited human feedback

SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

Abstract: Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.

Mon 28 SeptArtificial IntelligenceSound
The gist
Automatic systems usually rate speech quality with a single score, missing important details like emotion or timing. The authors created SpeechCritic, a method that learns to judge speech more finely by using just a small set of human comparisons. It improves an existing speech-language model by feeding it hints based on acoustic features that match human opinions. This approach works across different languages and makes the system better at explaining why one speech sample is better than another.
Open → 2609.34582v1

VoiceNet improves fine-grained understanding of voice emotions and styles

VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale

Abstract: Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acted speech. This paper introduces VoiceNet, a human-annotated representation-level benchmark for voice performance understanding on permissively-licensed in-the-wild speech. VoiceNet has two subsets: VoiceNet-Emo applies a 40-emotion taxonomy with three expert ratings per item, and VoiceNet-Ext, a preliminary subset, scores 57 talking-style attributes including speaking rate, vocal tension, breathiness, and register. The paper also releases Emolia, an emotion-annotated version of the Emilia corpus, with a curated rebalanced subset enriched by dense MOSS-Audio Thinking annotations. Two voice-text contrastive baselines train on this data: a 110M-parameter VoiceCLAP-Small for fast large-scale data filtering and a 7B VoiceCLAP-Large for state-of-the-art performance. Both outperform existing CLAP baselines, which sit near chance on VoiceNet-Emo. On VoiceNet-Emo, VoiceCLAP-Large aligns more closely with the aggregate expert consensus than individual experts agree with one another: a comparison against the majority label rather than evidence of surpassing human emotion perception. All systems evaluated here are voice-text embedding models: VoiceNet scores representation-level attribute recognition and retrieval, not end-to-end spoken-dialogue behaviour. Clustering and filtering uncurated speech corpora into subsets that span diverse talking styles and emotions remains an open challenge; VoiceCLAP embeddings offer a promising tool for this task. VoiceNet, Emolia, and VoiceCLAP are publicly available for research use.

Fri 25 SeptSoundArtificial Intelligence
The gist
Most existing systems can only recognize a few basic emotions from voices, often using acted speech that is less natural. The authors created VoiceNet, a new dataset that includes detailed human annotations for 40 emotions and over 50 voice style attributes in real-world speech recordings. They also developed VoiceCLAP models that learn to understand voice and text together, outperforming previous methods at identifying these fine-grained voice features. This work helps build tools that can better recognize subtle and complex voice expressions beyond simple emotion categories.
Open → 2609.32016v1

Speech data improves pinpointing topics in long transcripts

Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts

Abstract: Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.

Fri 18 SeptComputation and Language
The gist
Long audio transcripts can be hard to search because they contain a lot of extra information. This paper shows how using hidden speech features from the automatic speech recognition process can help find exactly where a topic is discussed in the transcript. The authors combine these speech features with text analysis to better locate relevant sentences without needing extra audio processing. Their method works best on structured talks like lectures or interviews but less well on casual conversations.
Open → 2609.21844v1

Spoken sarcasm detection relies more on words than voice tone

CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection

Abstract: Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under original, lexical-preserving, prosody-preserving, and approximately neutralised conditions. We evaluate handcrafted acoustic-feature systems, self-supervised learning (SSL) probes, and large audio language models (LALMs) on CMMA and MUStARD. For target-only Qwen3-Omni, lexical-preserving speech retains a 0.135--0.148 AUROC advantage over prosody-preserving speech after duration balancing, with cluster-bootstrap intervals above zero; alternative lexical resynthesis preserves this advantage. Acoustic interventions shift scores without consistently improving discrimination or changing binary predictions under the evaluated conditions. Context and interaction estimates vary across corpora. These findings distinguish acoustic sensitivity from sarcasm discrimination while exposing duration, identity, and transformation effects.

Tue 15 SeptSoundComputation and Language
The gist
Detecting sarcasm in spoken language can depend on the words people use, the way they say them (tone, speed), or both together. The authors created a way to test which of these clues speech recognition systems actually use to spot sarcasm. They found that the words themselves usually play a bigger role than the tone or prosody when identifying sarcasm in speech. This helps understand how sarcasm detection systems work and shows that changing tone alone doesn’t always change the system’s sarcasm judgment.
Open → 2609.16582v1