Summary
Knowing if someone wants to talk to a robot is something people do naturally by watching how others act. The researchers tested how well humans and different computer models can guess a person's intention by looking at either only body poses or full video from a robot's view. They found humans did slightly better when only body poses were shown, and much better when they could see the full video with the person highlighted. Even the best AI models that try to understand the situation with language and vision together did not match human intuition. This shows that while AI is improving, it still struggles to fully understand social cues as humans do.
human-robot interactionpose estimationegocentric videovision-language modelssocial intuitionintention predictionF1-ScoreHUI360 datasetmachine reasoninglightweight models
Authors
Raphael Lorenzo-Louis, Bertrand Luvison, Serena Ivaldi
Abstract
Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans, that relies on a combination of cues. We investigate how humans perform at predicting a person's intention to interact from a service robot's point of view, using pose-only or full video input, then benchmark different lightweight pose-based models and state-of-the-art vision-language models. We conducted our benchmark on the HUI360 dataset on a fixed pilot subset of 100 test tracks (25 positive, 75 negative). We found that with pose-only input, human annotators outperform lightweight trained pose models but not by large margins (+0.08 in F1-Score). But when given full egocentric video with a target bounding box, human annotators perform substantially better and largely outperform the Vision-Language Models (+0.2 in F1-Score). We also compared VLMs of different size and under different input conditions, and found that the best results do not correlate with model size. Our result confirms that predicting interactions is a challenging task for social robots and that reasoning-capable models are necessary but their actual reasoning capabilities alone do not suffice to match the social intuition of humans.