Papers for

policy modelers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Language models mimic farmer averages but miss individual decisions

The average-farmer illusion in language-model simulations of agricultural decisions

Abstract: Language-model agents are increasingly used as synthetic people in surveys and social simulations, yet their apparent realism is often judged from population averages or distributional similarity. We tested what such evidence actually establishes by comparing Claude, Codex and Kimi under four prespecified prompt designs with matched farmer decisions from China and four African countries. Some configurations reproduced observed means and adoption rates. However, their person-level predictions were weak; their decisions clustered around typical values and policy-relevant extremes were largely missing. Most strikingly, a simple generator fitted only to the observed marginal dis- tribution, and given no information about any farmer, achieved greater distributional similarity than every language-model configuration. Prompt additions produced conditional gains rather than uni- versal improvement: results varied with model, outcome, population and validation target. We call this the average-farmer illusion: a synthetic population can look realistic while failing to repro- duce who does what or how behaviour varies. We provide a claim-matched validation framework and reusable modular prompts that turn prompt construction into an auditable experimental process. Population-level resemblance should therefore be treated as the start of validation, not as evidence of individual simulation.

Mon 14 SeptArtificial IntelligenceComputation and Language
The gist
Language models are often used to simulate how people make choices, like farmers deciding to adopt farming practices. This paper shows that while these models can match overall averages seen in real farmers, they fail to predict what individual farmers actually do. In fact, a simple random model that only looks at overall data distribution can sometimes do better. The authors introduce a way to test models more carefully, so simulations can better reflect both group trends and individual behaviors.
Open 2609.15038v1