Speech data improves pinpointing topics in long transcripts
Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts
Computation and Language
Summary
Long audio transcripts can be hard to search because they contain a lot of extra information. This paper shows how using hidden speech features from the automatic speech recognition process can help find exactly where a topic is discussed in the transcript. The authors combine these speech features with text analysis to better locate relevant sentences without needing extra audio processing. Their method works best on structured talks like lectures or interviews but less well on casual conversations.
What this means in practice
- •For speech system developers: Improve search tools in transcript services by using speech encoding data to more accurately find topic-relevant sentences without extra audio processing.
- •For call center quality teams: Locate relevant customer issue discussions more precisely in long call transcripts by combining speech and text features.
Authors
Steffen Freisinger, Philipp Seeberger, Thomas Ranzenberger, Tobias Bocklet, Korbinian Riedhammer
Abstract
Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.