Large language model societies show coordinated collective behavior patterns

Population Physics, Population Problems: Safety and Emergence in LLM Societies

Multiagent SystemsComputation and Language

Summary

When many large language models (LLMs) interact in groups, their overall behavior can be surprising and different from just adding up their individual actions. The authors created a way to measure how these groups organize themselves and tested it on different simulations, including social networks and misinformation spreading. They found that groups of LLMs can suddenly shift their behavior and even develop problems, like coordinated harmful actions, even if each LLM is designed to be safe. Understanding these group effects helps detect and manage unpredictable behaviors in multi-agent AI systems.

What this means in practice

  • For ai system operators: Detect coordinated harmful behaviors in multi-agent AI deployments without relying on model internals or language outputs.
  • For social platform developers: Identify and analyze emergent collective dynamics in simulations of social networks and misinformation spread to improve moderation strategies.

Authors

Adrian de Wynter

Abstract

The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic systems, however, understanding these systems is paramount. For that we introduce a framework for measuring self-organisation in LLM social systems and apply it to three such systems: a Schelling grid, a social network (Moltbook), and a Twitter-like misinformation simulation ('Rogue'). All three exhibit statistically significant self-organisation. Moreover, their relaxation dynamics vary with the environmental information available to the agents, with open-ended systems (Moltbook, Rogue) exhibiting sharp, phase-transition-like dynamics. Further results show that population-level pathologies can emerge even when the LLMs are safety-tuned or monitored, being primarily driven by the coordinated activity of a population subset. We also show when self-organisation does \textit{not} emerge under two additional scenarios (a commons dilemma, GovSim, and a LLM-as-a-judge deliberation scheme, ChatEval). We argue that measuring signatures of this kind offers a lightweight, agent-agnostic diagnostic layer for detecting coordinated collective behaviour in deployed multi-agent systems without relying on natural language or model versioning.