Demographic synthetic panels may not simulate individuals despite matching survey totals

Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure

Computation and LanguageHuman-Computer Interaction

Summary

Synthesized survey panels aim to mimic real people’s answers to surveys, but this study finds that matching overall survey results doesn’t prove they really simulate individual respondents well. The authors compare different ways of generating and validating these panels and show that simply matching marginal responses can be explained by the way the data is generated, not by true individual-level simulation. In fact, sometimes just using overall population numbers works better than simulated panels. This means that current methods may not be reliable for creating synthetic individuals in surveys.

demographic synthetic survey panelsmarginal fidelityresponse contractpopulation prevalencemean absolute errorsimulation validationaggregate dataprobability elicitationlatent inclusion propensitiescheck-all responses

Authors

Alexander Doudkin

Abstract

Demographic synthetic survey panels are often validated by matching aggregate answers to published surveys. We test what that certificate establishes across six multiselect batteries from four survey organisations in three countries. The headline analysis is restricted to three instruments whose synthetic cohort and human target share the stated population frame; three other batteries remain sensitivity analyses. The response contract dominates measured fidelity. In the aligned instruments, committed sets leave 66 of 128 model-battery option slots empty in panels of up to 500 respondents, versus 0 of 128 under per-option probability elicitation. Across eight uncapped model-instrument comparisons, probabilities reduce option-marginal MAE by 4.53 to 7.30 points. The capped instrument reverses on two models until the vectors are projected onto its stated maximum. These are measurement effects: human targets are realised check-all responses, whereas the vectors are latent inclusion propensities. Published marginal agreement also fails to discriminate respondent simulation from direct population estimation. On nine aligned model-battery pairs, a no-persona population-prevalence query averages 6.27 MAE versus 12.39 for committed panels and wins all nine comparisons. Constraint-aware probability vectors average 5.34 and beat the query on four of nine, so the baseline challenges the validation criterion rather than proving direct estimation uniformly best. On three unpublished demographic cells, neither approach beats reciting the national distribution. Population-marginal agreement is therefore evidence about an elicitation contract and an estimand obtainable without simulated respondents, not evidence of individual simulation.