Papers for

customer service automation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

RoboCafé improves robot interactions over many days in public spaces

RoboCafé in the Open: Interaction Continuity in Long-Term Public Human-Robot Interaction

Abstract: As robots remain in public spaces over extended periods, they must maintain interaction continuity by preserving and correctly applying context as people, encounters, and circumstances change. To study interaction continuity in long-term public human-robot interactions, we developed RoboCafé, an autonomous conversational coffee robot designed to support repeated interactions through task-aware dialogue, real-time multimodal perception, and memory of prior encounters. We deployed RoboCafé for 12 days in a university building, where it received 148 orders. The deployment involved repeat customers, passersby, changing groups, and back-to-back orders that repeatedly crossed the boundaries assumed by the system's order-centered interaction model. We found that successful interaction continuity requires a robot to determine who is currently present, which prior context belongs to whom, where interactions begin and end, and whether its representation of an interaction matches what is occurring in the physical world. From these observations, we derive four system design requirements for maintaining interaction continuity in longitudinal public human-robot interactions: contextual interaction state, persistent person grounding, explicit interaction life-cycle management, and interaction observability.

Wed 23 SeptRobotics
The gist
Robots in public places need to remember who they are talking to and what happened before in order to have smooth conversations over time. The authors created RoboCafé, a coffee robot that talks to people, remembers past orders, and notices who is nearby. They tested RoboCafé for 12 days at a university and found it had to carefully figure out who was present and which conversation was happening to keep interactions going well. Based on this, the authors describe four important features robots need for long-lasting public talks.
Open → 2609.27475v1

Agentic AI workflows combined to improve task accuracy and cost control

Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost. A natural deployment policy is to use the workflow with the highest average performance, but this can be suboptimal because different workflows may succeed on different instances. We study a portfolio-and-selector paradigm in which a firm runs multiple workflow executions and selects the final answer after observing their outputs. Additional executions may uncover correct answers that the best standalone workflow misses, but they consume compute and introduce plausible distractors that complicate final selection. We formulate this as a workflow portfolio problem in which the firm jointly chooses run size and allocation across workflow types. We summarize selector quality through an odds-lift index and derive sharp bounds on the value of workflow variety. For finite workflow pools, we develop exact formulations, linear programming relaxations, randomized rounding procedures, and computable performance certificates. For large implicit workflow classes, we derive a finite-dimensional dual and an ellipsoid method using a pricing oracle to identify workflows with high weighted accuracy net of recurring compute cost. Under a weak condition, the method obtains a near-optimal solution to the relaxation with polynomially many oracle calls. We evaluate the framework on three datasets: ABCD, Schema-Guided Dialogue, and HotpotQA. Relative to the best standalone workflow, portfolio optimization improves held-out selector accuracy by 3.1, 7.5, and 0.9 percentage points, respectively. Dual-guided workflow generation adds 3.5 points on ABCD and 24.1 on HotpotQA, with no additional gain on Schema-Guided Dialogue.

Wed 16 SeptArtificial Intelligence
The gist
When AI systems try to solve problems, they often use different methods that have different strengths, costs, and ways of checking answers. The paper studies how running multiple AI methods together and then choosing the best result can work better than just picking one method all the time. The authors create a mathematical way to decide how many times to run each method and how to pick the best final answer while considering computing costs. They tested this on three different tasks and showed that combining methods improved accuracy by a few percentage points compared to using just the single best method.
Open → 2609.18126v1

Pace reduces first response delay in dialogue systems under load

PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving

Abstract: We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a QoE objective and minimizes it under quality/cost constraints. Unlike prior work on cascaded routing, semantic caching, or adaptive retrieval, PACE jointly controls which answer source composes the response and what fills the waiting window. Deployed on a humanoid-robot sales service, it combines three mechanisms: a load-adaptive cascading router, a joint path-filler controller, and volatility-aware cache admission. On 75k CarQA requests, the cascade halves pure-LLM PTFR at P95 (0.29 vs 0.53s at c16). The adaptive controller reaches 0.41s P95, outperforming RAG by 2.4 times at high load with equal quality. The filler controller cuts calls by 94% with zero conflict. Volatility-aware admission reduces stale answers from 86% to 0%. A gating rule ensures the controller never worse than the baseline, with exposure bounded by one hold period. This is the first quantification of filler-answer conflict risk in deployed services.

Wed 9 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceRobotics
The gist
Waiting for a reply from a chatbot or robotic assistant can feel slow, especially when many people are using it at once. The authors created a system called PACE that smartly chooses where answers come from and what to show while waiting, to make the first reply feel faster. They tested PACE on a robot helping customers with car questions and found it cut delays nearly in half and avoided stale or conflicting answers. This approach balances quickness with answer quality and cost.
Open → 2609.10372v1

Hierarchical bayesian model improves robot group talk and clarity

HiBRIDGE: A Hierarchical Bayesian Neural Network Framework for Interpretable Dialogue Management in Group-Robot Interaction

Abstract: In multi-party human-robot interaction, a robot must continuously decide whom to address and what to say to participate effectively in the conversation. In real-world interactions, this is challenging because several behaviours may be plausible at the same time: a robot might continue a topic with one participant, involve another through a question, or address the whole group, with the appropriate choice depending on both whom it addresses and the interaction context. Current approaches remain limited in representing uncertainty when several behaviours are plausible and in structuring decisions into semantically meaningful intermediate steps that make robot decisions easier to interpret. Addressing these, we present HiBRIDGE, a hierarchical Bayesian neural network framework for group-robot dialogue management. Its Bayesian formulation enables uncertainty-aware prediction and robust learning from limited interaction data, while the hierarchical approach formulates behaviour selection as a structured, multi-stage decision process. We further use decision-tree surrogates to investigate whether this structure can support more interpretable explanations. Across three offline group-HRI datasets, our findings show that Bayesian formulations outperform their deterministic counterparts and several state-of-the-art baselines. Next, through an online study (N=20), we show that explanations derived from the hierarchical model are rated as more helpful for understanding robot behaviour and are preferred over those derived from the flat model. Finally, through our in-person study (N=12), we demonstrate the feasibility of HiBRIDGE for autonomous real-time group interaction, with both hierarchical and flat Bayesian variants positively perceived. Overall, HiBRIDGE combines strong predictive performance with a structured decision process that supports more interpretable explanations of robot behaviour.

Tue 8 SeptRobotics
The gist
In conversations with groups, a robot must decide who to talk to and what to say, which can be tricky when multiple choices seem right. The authors created HiBRIDGE, a system that uses a step-by-step approach and probability to handle uncertainty and learn from little data. This helps the robot make decisions that are easier to understand and explain to humans. Tests show their method outperforms others and that people find its explanations clearer and more useful during real interactions.
Open → 2609.08678v1