Large language models show varied political bias based on language and prompts

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

Computation and LanguageComputers and Society

Summary

Large language models, which help provide information, can have different political biases depending on how questions are asked and in which language. The researchers created a thorough test to measure these biases more reliably by changing many factors like wording and answer styles. They found most models tended to lean Libertarian-Left but results shifted with different instructions or languages. Smaller models sometimes seemed neutral because their answers clustered around the center, which might not mean true neutrality. The way these models are prompted can also affect how they behave in tasks like detecting hate speech or analyzing sentiment, but these effects depend on the task and data used.

Large Language ModelsPolitical biasPromptingCross-lingual differencesPersona promptingQuantizationHate speech detectionSentiment analysisMeasurement artifactsPolitical compass

Authors

Luka Debevc, Nishan Chatterjee, Antoine Doucet, Senja Pollak, Matej Martinc

Abstract

Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.