Predicting the scale limits of social mechanisms in agent societies

2026-08-24Multiagent Systems

Multiagent SystemsSocial and Information Networks
AI summary

The authors studied how groups of AI language agents behave when the group size changes, especially when it gets very large. They created a test, called an audit, that predicts if certain social behaviors like sharing gossip or punishing others still work as more agents join. Their experiments showed that simple factors can determine if these behaviors last or fade away with bigger groups. They also found that how information is presented affects how agents react, and their predictions mostly worked across different AI setups. This method helps decide which social rules in AI communities can be reliably used no matter how many agents are involved.

language-model agentscollective behaviourreciprocityconsensuspunishmentgossipsocial mechanismsscalingauditpopulation dynamics
Authors
Zengqing Wu, Chuan Xiao
Abstract
Societies of interacting language-model agents offer a controllable and repeatable way to study collective behaviour at scales that would be difficult to test with people. Their scientific value, however, depends on whether a social mechanism that works in a small group still operates when thousands of agents interact, and testing this directly requires costly large-scale runs. Here we introduce an audit that predicts a mechanism's fate as a population grows. It asks how often the mechanism can act, whether agents use the information it supplies, and whether the measurement itself creates apparent scale effects. Controlled experiments show that a single structural term can decide whether reciprocity, consensus or punishment survives scaling. For gossip, the population at which the mechanism fails is set by the reach and lifetime of its messages. In language-model societies, agents respond not only to social information but to how it is expressed: counts and percentages led to different scale behaviour. Predictions made before execution held on third-party code and a second model family, while a failed prediction exposed the boundary of the finding. The audit provides a prospective way to decide which social mechanisms can be interpreted across population scales.