Benchmark compares human and AI social choices for companion robots
InterSocialBench: Benchmarking Human and LLM Preferences for Companion-Robot Social Behavior
Robotics
Summary
Robots that keep people company face many everyday situations where there isn’t just one right way to act. People have different preferences for how robots should behave in those moments. The authors created a big set of scenarios and asked humans to pick their favorite robot behaviors, then compared those to responses generated by large language models simulating different personalities. They found that humans showed more diverse preferences and less agreement than the AI models. Their benchmark helps measure how well AI can match human social behavior choices without forcing a single answer.
What this means in practice
- •For companion robot developers: Compare large language model responses to human social preferences to improve robot behavior selection.
- •For human-robot interaction designers: Evaluate the diversity and acceptance of social behaviors in robots across different scenarios and personalities.
Authors
Yaodan Xu, Boyang Guo, Yuqing Gu, Qingxin Zhang, Yiwen Deng, Meng Liu, Lintian Li
Abstract
Companion robots face everyday situations in which several feasible behaviors may be appropriate, yet different people prefer different responses. We introduce InterSocialBench, a benchmark of 210 domestic scenarios and 18 high-level behaviors, pairing judgments from 100 human participants with 23,520 responses from seven large language models under 16 personality conditions. Each human annotation preserves a preferred action alongside explicitly appropriate and inappropriate candidates. A structured construction pipeline covers behavioral alternatives, competing situational cues, and relevant history and future tasks. Evaluation distinguishes preferred-choice agreement from explicit rejection, using scenario-grouped splits for trainable predictors. Simple frequency and persona-voting baselines illustrate these objectives. Across the tested prompts, model and human behavior distributions differ, and the diversity gap remains after matching response counts: humans exhibit 4.68 distinct choices per scenario, compared with 2.06--3.46 for the models. Human scenario-level plurality agreement is 51.5%, describing disagreement rather than a universal prediction ceiling. InterSocialBench supports evaluating social behavior selection without replacing individual judgments with a single consensus label.