Human-robot communication improves by measuring listener knowledge gaps
Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration
RoboticsHuman-Computer Interaction
Summary
Robots sometimes fail and need help from people nearby, but it's hard to know how to ask for help in a way everyone understands. The authors created a game and a dataset to test how well different robot requests work when people have different knowledge levels. They found that robot-generated requests help beginners somewhat but still leave a big gap compared to experts. Their work helps scientists measure and improve how robots and people talk when working together.
What this means in practice
- •For robotic system designers: Test and improve robot failure recovery communication by measuring how people with different knowledge understand robot requests.
- •For virtual assistant developers: Design more effective help requests by accounting for differences in user knowledge when generating language using large language models.
Authors
Promise Ekpo, Teju Vijay, Dhruv Mandalik, Tisha Jain, Arman Ibrayeva, Sunishka Sil, Stefanie A. Tellex, Angelique Taylor
Abstract
Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We introduce Listener Differences in Human-Robot Interaction (LD-HRI), a game, dataset, and benchmark that evaluates speakers through human listener performance. Our evaluation examines request properties, large language model (LLM) speakers, and inverse-semantics request-selection algorithms under controlled differences in listener information. The corpus contains 446 human-written requests and 1{,}302 listener trials. We additionally evaluated 24 frozen LLM-written requests with 70 human listeners across 560 trials. Novice success is descriptively higher with model-written requests across all four tasks, yet both request sources leave substantial expert--novice gaps, including 16 percentage points for LLM requests. LD-HRI makes these gaps measurable, providing a foundation for designing more robust communication in human-robot and human-agent interaction.