Simulated agents struggle to reflect real public opinions in debates

From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction

Artificial Intelligence

Summary

The paper looks at using computer-created characters to mimic how real people discuss important issues. The authors found that these virtual characters often don’t match real population opinions well, and sometimes they wrongly reverse who agrees or disagrees by group. Even though the characters make good arguments and change opinions during discussions, much of this change happens without real interaction with others. The study shows that making virtual debates realistic is complicated, especially in representing diverse viewpoints, but the method can still help find different arguments.

large language modelsmulti-agent systemspublic deliberationpersona agentsopinion dynamicspopulation representationpolicy simulationargument generationcensus datastance movement

Authors

Chaemin Jang, Junsik Min, Jaewoo Choi, Donggyu Lee, Haiin Lee, Junyoung Park, Namhee Kim, Hyunwoo Kim, Jungwon Kim, Juho Kim, Nuri Kim, Jihee Kim

Abstract

Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such simulations to be informative, persona agents should reflect population opinion patterns and interaction should shape their conclusions. We evaluate whether LLM-based deliberation can meet these two conditions using census-grounded Korean personas debating real policy questions benchmarked against national surveys. Persona agents do not reliably reproduce population opinion patterns: responses are often far more concentrated and frequently reverse demographic differences in the human data. Deliberations nonetheless produce reasoned, reciprocal, and varied arguments alongside substantial stance movement. Yet much of this movement does not require peer exchange: sealed-monologue agents change position at similar rates and reach nearly the same final balance as full debates, while groups initialized with very different positions often converge to similar endpoints. Anchoring population-informed starting positions, meanwhile, sharply suppresses updating. Thus, population representation, argument generation, and interaction-driven opinion change do not necessarily go together. The simulations readily surface arguments on both sides, though whether they capture the diversity of human perspectives remains untested, leaving open a promising role for argument surfacing even as population simulation requires further validation.