INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval
2026-08-17 • Sound
SoundComputation and Language
AI summaryⓘ
The authors created INSPIRE, a new test to see how well speech search systems can follow different natural language instructions to find specific sounds or speakers. They tested four types of systems and found that none could handle all kinds of instructions well. Systems using text understand meaning better but struggle with voice characteristics, while those using speech handle sounds better but have trouble following detailed instructions. The authors suggest developing new systems that combine both strengths to understand instructions better.
speech retrievalnatural language instructionsaudio-language modelsself-supervised learningcontrastive learningsemantic contentparalinguistic featuresbenchmark dataset
Authors
Chen-An Li, Hung-yi Lee
Abstract
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.