Papers for

call center analytics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Cue grounded method improves spoken topic segmentation with zero training

Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents

Abstract: Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.

Mon 28 SeptComputation and LanguageArtificial Intelligence
The gist
Breaking long spoken documents into topics helps people find and understand information better, but it's hard to do this well when topics vary in size. The authors designed a new way that doesn't require teaching the computer beforehand. Their method looks for special words or phrases that usually signal a new topic, and if those aren't clear, it uses the overall meaning to guess where topics change. This method works better than others, even with imperfect speech-to-text transcripts, and is cheaper to run with big language models.
Open → 2609.34425v1

Speech language models improved for emotion recognition with linear classifier

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

Abstract: SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.

Thu 17 SeptComputation and LanguageArtificial IntelligenceSound
The gist
Detecting emotions in speech is tricky because existing models often guess wrong or use labels not in the expected set. The authors changed how a speech language model reads emotions by adding a simple layer that picks the emotion directly from internal data without changing the main model. This approach reduces mistakes and works better, especially when using text from automatic speech recognition, revealing how subtle language biases connect with emotions. It also helps understand which words the model links to each emotion.
Open → 2609.20081v1