Language model persona prompts change output without fixing bias

The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs

Computation and Language

Summary

The paper investigates whether giving large language models (LLMs) different personas through prompts can reduce their biases. The authors found that while changing the persona does affect what the model says, it doesn't change the model's deeper internal biases. Instead, these persona changes mostly influence the final output without altering the model's underlying associations or how it processes information. This means that prompt-based methods to steer models might only mask biases rather than truly reducing them.

large language modelsbiaspromptingpersona conditioninginstruction tuningoutput modulationinternal structureassociative biasclosed-form question answeringinter-trait covariance

Authors

Ziyue Feng, Hongbo Fang, James A. Evans

Abstract

Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.