LLM digital twins may not reduce human data collection as expected
When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability
Artificial Intelligence
Summary
Collecting human data repeatedly is time-consuming, so computer models called digital twins try to create personalized responses to reduce this need. The authors study whether these digital twins can actually replace human measurements while keeping results trustworthy. They find digital twins can mimic average group behavior well but struggle to identify individual differences accurately. Even with better models and extra information, these improvements don’t always mean less human data is needed. This work shows it’s more important to check if AI models support valid conclusions rather than just copying human data patterns.
LLM (Large Language Model)digital twinstatistical substitutabilitybehavioral fidelityinferential criterionhuman measurementmixed-subject inferenceprediction-powered inferencescientific inference
Authors
Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph
Abstract
LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.