LLM batching method cuts survey costs and improves accuracy

SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys

Computation and LanguageComputers and Society

Summary

Large language models can mimic how people answer surveys, but asking one question at a time wastes time and energy. The authors created a method called SCBO that groups related questions together, shares useful examples, and asks easier questions before harder ones. This makes the process faster, uses fewer computer resources, and helps the model give better answers. They tested SCBO with different surveys and model types and saw clear benefits.

What this means in practice

  • For social science survey teams: Run large-scale simulated surveys more efficiently by grouping questions and sharing response examples within language model prompts.
  • For market research companies: Improve accuracy and cut response costs when using AI-driven simulations of consumer opinions across multiple survey questions at once.

Authors

Yuanzi Li, Lingjie Wang, Zihang Tian, Lei Wang, Xu Chen

Abstract

Large Language Models (LLMs) offer a scalable way to simulate survey respondents using demographic profiles and observed reference responses. However, the conventional approach of predicting one question per prompt repeatedly encodes the same context, limits each target to a narrow set of reference responses, and prevents later predictions from using information in earlier answers. Predicting multiple questions in one prompt can reduce these costs, share a broader pool of references, and let later predictions build on earlier ones. This requires forming coherent batches, selecting shared references, and ordering questions and references effectively. We propose Semantically Coherent Batching and Ordering (SCBO), a training-free framework that addresses these challenges. SCBO first uses an LLM to extract compact semantic representations from survey items and filter out template noise. It then groups related questions into batches and builds a shared reference bank using target-specific retrieval and centroid-based completion. Finally, it orders target questions from easy to hard and arranges references according to their semantic alignment with those questions. Experiments on four large-scale survey datasets and four LLMs show that SCBO substantially reduces token consumption and inference time while generally improving prediction accuracy over a non-batched baseline. Code is available at https://anonymous.4open.science/r/SCBO-41D8.