Bayesian method improves robot questions to understand human instructions
Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning
RoboticsArtificial IntelligenceHuman-Computer Interaction
Summary
Robots often have trouble understanding unclear or incomplete human commands. The authors propose a way for robots to ask the most helpful questions to figure out what humans really mean. They use a combination of a statistical approach called Bayesian optimization and large language models to guess possible tasks and choose questions that best reduce confusion. This approach helps robots understand instructions better while asking fewer questions. It works well in both simulations and real robot tasks.
What this means in practice
- •For robotics engineers: Enable robots to efficiently clarify ambiguous human commands and generate precise task plans.
- •For voice assistant developers: Improve interactive systems to ask better clarification questions for ambiguous or incomplete spoken instructions.
Authors
Huao Li, Carson Sobolewski, Augustinos Saravanos, William Tan, John Karigiannis, Chuchu Fan
Abstract
Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.