Robot learns to help find objects using language and actions

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

Robotics

Summary

Helping robots assist people at home, especially older adults or those with disabilities, is hard because robots must understand what humans want through speech and body movements. The authors show a way for robots to learn how to plan their actions automatically by practicing in a virtual setup that mimics real human behavior. Their method lets robots decide the best thing to do next without people programming every step. When tested with real users, their approach helped robots work well and made tasks easier to complete.

What this means in practice

  • For robotics developers: Create robot assistants that learn effective ways to help humans find objects at home by combining speech and physical cues.
  • For smart home device makers: Develop home robots that adaptively assist users with locating items using natural communication and actions.$Commercial implications: Enables consumer home robots that provide task support through natural interaction, improving usability and customer satisfaction.

Authors

Afagh Mehri Shervedani, Siyu Li, Natawut Monaikul, Bahareh Abbasi, Barbara Di Eugenio, Miloš Žefran

Abstract

Robot assistants for older adults and people with disabilities need to perform collaborative tasks with users effectively. The core component of these systems is an interaction manager whose job is to observe and assess the task and infer the state of the human and their intent for the robot to choose the best course of action. Due to the sparseness of the data in this domain, the policy for such multimodal systems is often crafted by hand; as the complexity of interactions grows, this process is not scalable. This paper proposes a reinforcement learning (RL) approach to automatically generate the multimodal policy of the robot. Our system focuses on a realistic scenario where a robot assists a user in locating objects within a home environment, managing multimodal signals, including language and physical actions, to select the best action. In contrast to traditional dialog systems, our agent is trained with a simulator that uses human data and can deal with multiple modalities. We use a simple high-level reward function that needs no fine-tuning and enforce some preconditions to speed up the training process. A human study evaluating the system in a real-world setting demonstrates promising results, indicating high usability and effective task completion. This RL-based approach offers a scalable and interpretable alternative for designing interaction managers in multimodal human-robot collaborations.