Language models tested for self and social awareness skills

AwarenessBench: Assessing Cognitive Capabilities of Language Models

Computation and LanguageArtificial Intelligence

Summary

Evaluating how well language models understand themselves and social situations is important as they behave more like conscious beings. The authors created AwarenessBench, a large test with many tasks covering self-knowledge, awareness of others, and thinking about thinking. They tested advanced language models and found that while some beat humans overall, most still struggle with knowing their own limits and feelings. They also showed that being good at language or reasoning doesn’t mean a model is better at awareness.

What this means in practice

  • For ai developers: Use AwarenessBench to evaluate and improve the cognitive and awareness skills of language models they build for better interactive AI systems.
  • For chatbot designers: Assess chatbot abilities in self-reflection and social understanding to create more relatable and context-aware conversational agents.

Authors

Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang, Shuo Chen, Qiner Lyu, Briana Zhang, Peiran Yang, Kyle Xue Chen, Haoyuan Shi, Yu Wang, Wei Xu

Abstract

As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social awareness, and situational awareness, covering 15 cognitive functions and 14,381 samples. Evaluating 18 state-of-the-art LMs, we find that all consistently surpass random baselines, with more advanced models performing better. We further compare LMs with human performance across three demographic groups, where the best-performing model surpasses human averages overall, but most still fall markedly short in metacognition and self-awareness. Finally, we show that awareness is a distinct capability: progress in language modeling or reasoning does not necessarily translate into improved cognition.