Large language models simulate intersectional synthetic identities with a budget of one to two dimensions

2026-08-24Computers and Society

Computers and Society
AI summary

The authors studied how well large language models (LLMs) can mimic people's opinions when given multiple combined identities, like race and religion together. They found that real people's views become more unique when multiple identities intersect, but LLMs mostly respond based on just one identity at a time. Trying different ways to prompt the models didn't fix this issue, and the models often ignored important identity factors like race and religion. This means LLMs do not truly capture the complexity of intersectional identities in surveys.

Large Language ModelsIntersectionalitySynthetic RespondentsDemographic PersonasAmerican Trends PanelSurvey SimulationPrompting StrategiesAdditive ModelsRaceReligion
Authors
Virgile Rennard, Christos Xypolopoulos
Abstract
Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.