Pace reduces first response delay in dialogue systems under load

PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving

Computer Vision and Pattern RecognitionArtificial IntelligenceRobotics

Summary

Waiting for a reply from a chatbot or robotic assistant can feel slow, especially when many people are using it at once. The authors created a system called PACE that smartly chooses where answers come from and what to show while waiting, to make the first reply feel faster. They tested PACE on a robot helping customers with car questions and found it cut delays nearly in half and avoided stale or conflicting answers. This approach balances quickness with answer quality and cost.

What this means in practice

  • For chatbot platform engineers: Improve chatbot responsiveness under heavy use by adaptively routing answer sources and controlling filler content to reduce user-perceived delay.
  • For customer service automation teams: Deploy on conversational robots to cut wait time and minimize stale or conflicting answers for real-time customer support.$Commercial implications: Enables robotic sales assistants to deliver faster, more reliable customer interactions, boosting service quality and satisfaction.

Authors

Lin Huang, Yujuan Tan, Weisheng Li, Lixiang Zeng, Kun Yang, Suihan Xiao

Abstract

We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a QoE objective and minimizes it under quality/cost constraints. Unlike prior work on cascaded routing, semantic caching, or adaptive retrieval, PACE jointly controls which answer source composes the response and what fills the waiting window. Deployed on a humanoid-robot sales service, it combines three mechanisms: a load-adaptive cascading router, a joint path-filler controller, and volatility-aware cache admission. On 75k CarQA requests, the cascade halves pure-LLM PTFR at P95 (0.29 vs 0.53s at c16). The adaptive controller reaches 0.41s P95, outperforming RAG by 2.4 times at high load with equal quality. The filler controller cuts calls by 94% with zero conflict. Volatility-aware admission reduces stale answers from 86% to 0%. A gating rule ensures the controller never worse than the baseline, with exposure bounded by one hold period. This is the first quantification of filler-answer conflict risk in deployed services.