Adaptive model and history choices improve social survey simulations
Adaptive Resource Allocation for Effective and Efficient LLM Social Survey Simulation
Computers and Society
Summary
Simulating social surveys with large language models (LLMs) can be slow and costly when using the same strong model and lots of past answers for every question. The authors found that choosing the best language model and the right amount of past answer history for each survey question can give better results at lower cost. They created a system called E2Sim that predicts which model and amount of history to use for each question based on the respondent's persona and previous answers. Testing on real survey datasets showed their method improves accuracy and reduces computing resources compared to fixed strategies.
What this means in practice
- •For data science teams: Optimize simulation of social surveys by dynamically picking models and history length to improve accuracy and lower computational cost.
- •For marketing analytics groups: Create more efficient customer survey simulations by selecting appropriate AI models and relevant past responses per query.
Authors
Yuanzi Li, Xueyang Feng, Junhao Wang, Lei Wang, Xu Chen
Abstract
Large Language Models (LLMs) enable scalable social survey simulation, yet existing pipelines typically use the same strong general-purpose model and a fixed, often large, respondent history for every respondent-question request. This uniform approach overlooks three factors. First, stronger models may rely on their own knowledge rather than respondent-specific evidence while costing more. Second, additional history can help when evidence is limited, but irrelevant responses may add noise and increase input length. Third, the preferred model and history budget can depend on each other. We propose E2Sim, an adaptive resource allocation framework that jointly selects a model and history budget for each request. Given a respondent persona, ranked response history, and target question, a lightweight policy predicts the accuracy and cost of each configuration and selects the most suitable one. We use respondent-history drop-and-swap augmentation to improve robustness to incomplete or variable histories, and a margin-based curriculum that progresses from easier allocation decisions to harder ones. Experiments on four real-world social survey datasets, multiple model pools, and different history-budget spaces show improvements over the oracle best-fixed configurations, with accuracy gains of up to 6.3 percentage points and cost reductions of up to 55.3 percent. Code is available at https://anonymous.4open.science/r/E2Sim-DE66.