Large language models generate coherent survey responses across questions
Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models
Artificial Intelligence
Summary
Surveys often ask the same person many related questions, so their answers should make sense together. Existing methods simulate answers for single questions well but fail to keep answers consistent across an entire survey. The authors created a new method that trains two language models: one for each question's answer patterns and another for how answers relate across questions for the same person. Their method better mimics how real people respond to whole surveys and works well even on new populations or questions.
What this means in practice
- •For market researchers: Use simulated complete survey responses to test pricing and stocking strategies before real data collection.$Commercial implications: Enables market research firms to sell more effective survey simulations that improve decision-making and profitability.
- •For survey methodologists: Generate realistic synthetic survey data with coherent answers across all questions for testing survey designs and analyses.
Authors
Ji Huang, Mengfei Li, Shuai Shao
Abstract
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.