Social science reveals how large language models affect people and society

Mapping the Emerging Social Science of Large Language Models

Computers and SocietyArtificial IntelligenceComputation and Language

Summary

Large language models (LLMs) are powerful tools that influence how we communicate, learn, work, and create. The authors studied many research papers to organize social science studies about LLMs into three big groups: how the models behave like social minds, how multiple models interact like societies, and how humans use and are affected by LLMs. They found many topics in these groups, such as bias, trust, creativity, and education. This organization helps researchers better understand the social effects of LLM technology.

large language modelssocial scienceK-means clusteringlatent Dirichlet allocationstructural topic modelingmodel behaviorcollective intelligencehuman-computer interactionbiastrust

Authors

Yi Yang, Xiao Jia, Zeyun Dong, Chenzhang Wang, Zhanzhan Zhao

Abstract

Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.