Large language models struggle to understand hidden user needs in real life

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

Artificial Intelligence

Summary

Many AIs that talk like humans can help with daily problems, but they usually follow exact instructions and miss what users don’t directly say. The authors created a big test called xDailyBench to see how well these models handle real-world questions where people don’t explain everything. They found that while models do okay with clear requests, they often fail to guess what users implicitly want. This shows that current AIs need to get better at understanding people’s unspoken needs to be truly useful in everyday life.

large language modelsbenchmarkimplicit requirementsexplicit instructionsreal-world taskstask evaluationuser contextAI assistancenatural language understanding

Authors

Yongchang Peng, Qingshui Gu, Liya Zhu, Ge Zhang, Duo Wang, Haodong Wang, Jingzhe Ding, Tianhao Yu, Letian Gao, Yongjie Zhong, Chaoxin Li, Zixin Su, Jinchao Tao, Xingyu Ma, Xin'ao Guo, Feng Tian, Shiyuan Dong, Xiaoyan He, Sen Liu, Xin Chen, Jiajun Li, Zejia Zhang, Xi Lin, Wen Zhang, Yi Zhu, Duju Zeng, Xiang Gao, Yunyang Wang, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

Abstract

Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.