Benchmarking large language model agent societies against human behavioural distributions

Computation and Language

Summary

The authors introduce SILICA, a tool designed to test if large language model (LLM) agents behave like humans in social experiments, if results hold when experiment details change, and if social behaviors are real or just copied from training data. They tested 12 models across 5 environments with real human data for comparison. Results showed models only matched humans at the start of experiments, failed to replicate long-term cooperation, and were highly sensitive to small changes like action order. Only one model showed reasoning aligned with incentives, and social conventions largely formed from pre-learned patterns rather than genuine negotiation. The authors conclude current LLM-based societies are useful for exploration but not yet reliable for solid conclusions.

Authors

Raad Bin Tareaf

Abstract

Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.