Language model persona prompts change output without fixing bias
The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
Computation and Language
Summary
The paper investigates whether giving large language models (LLMs) different personas through prompts can reduce their biases. The authors found that while changing the persona does affect what the model says, it doesn't change the model's deeper internal biases. Instead, these persona changes mostly influence the final output without altering the model's underlying associations or how it processes information. This means that prompt-based methods to steer models might only mask biases rather than truly reducing them.
large language modelsbiaspromptingpersona conditioninginstruction tuningoutput modulationinternal structureassociative biasclosed-form question answeringinter-trait covariance
Authors
Ziyue Feng, Hongbo Fang, James A. Evans
Abstract
Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.