Calibrating social simulators against real market network data
Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
Artificial IntelligenceMultiagent SystemsSocial and Information Networks
Summary
Many social simulators that try to mimic real-world social behaviors stop at just looking similar without checking how accurate their parameters are. The authors present a new detailed method that tests and adjusts these simulators using uncertainties and diagnostic checks. They apply this to a second-hand luxury resale market, showing that while some behaviors in the model can be recovered, key parts of the real market aren’t fully captured by simple models without interactions. Their approach finds where the models miss and helps guide improvements.
What this means in practice
- •For social data scientists: Improve social simulation models by systematically testing and calibrating them against observed network features with quantified uncertainty.
- •For market analysts: Better assess and refine predictive models of consumer behavior in resale markets by identifying which network features models fail to reproduce.
Tested on one dataset.
Authors
Tengfei Shao, Chao Li, Xu Wang, Masayuki Goto
Abstract
Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.