Cue grounded method improves spoken topic segmentation with zero training

Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents

Computation and LanguageArtificial Intelligence

Summary

Breaking long spoken documents into topics helps people find and understand information better, but it's hard to do this well when topics vary in size. The authors designed a new way that doesn't require teaching the computer beforehand. Their method looks for special words or phrases that usually signal a new topic, and if those aren't clear, it uses the overall meaning to guess where topics change. This method works better than others, even with imperfect speech-to-text transcripts, and is cheaper to run with big language models.

What this means in practice

  • For podcast platform teams: Segment long spoken episodes into meaningful topics enabling easier navigation and summarization without additional training data.
  • For call center analytics teams: Automatically identify topic shifts in recorded calls to improve analysis and customer support quality using robust cue-based methods.

Authors

Suhwan Choi, Myeongho Jeon, Myungjoo Kang

Abstract

Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.