Voice dialogue system plays pre-recorded lines faster with low latency
RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue
Sound
Summary
Many voice assistants struggle to replay exact, pre-recorded responses quickly during conversations. The authors designed RePlay, a system that retrieves and plays back prerecorded lines with faster response times in multi-turn dialogues. They found a way to predict when a reply is ready to be retrieved early in the model’s processing, so it can quickly fetch the correct audio snippet. In tests, RePlay responded up to seven times faster than traditional systems but traded some accuracy in delivering the exact prerecorded line. Users generally preferred its faster, smoother interaction.
What this means in practice
- •For voice assistant developers: Provide faster multi-turn voice responses by retrieving prerecorded lines to reduce user wait times during interactions.
- •For call center technology teams: Implement low-latency response playback using prerecorded lines in interactive voice response systems to improve customer experience.
Authors
Sathvik Udupa, Naveen Kumar, Ryan Folmsbee
Abstract
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and playing pre-recorded lines. Using probing, we identify the layer and frame at which the upcoming response becomes recoverable, and use this hidden state as the retrieval query. RePlay retains only the layers up to that point and replaces text and speech generation with lightweight turn-taking and retrieval heads. In simulated multi-turn interviews, RePlay reaches a median latency of 383 ms, 3 to 7 times lower than ASR-LLM cascades of comparable dialogue quality, at the cost of lower exact-line accuracy. In a user study, participants preferred RePlay in 63% of ratings versus 12% for a fast cascade with a small LLM (p = 0.008), and showed a non-significant preference (46% vs. 21%) over a slower cascade with a stronger LLM.