Summary
Language models are increasingly used to access political information, but what they refuse to talk about can reflect censorship styles from authoritarian governments. The researchers found that models developed under such regimes tend to avoid helping people organize protests, especially about their own country, reflecting the state’s political fears rather than universal ideas of harm. These models are less strict when the same prompts concern other countries and can be fooled by slight changes in wording, making their censorship uneven. Models from Western countries show stronger resistance to this kind of manipulation. Overall, the way models filter content resembles old-fashioned censorship but is less flexible and more blunt.
language modelsinformation controlauthoritarian censorshipguardrailspolitical informationcontent moderationadversarial paraphrasecollective actionAI biasstate influence
Authors
Menglin Liu, Yao Yu, Tong Wu, Chunran Zhang, Ge Shi
Abstract
As large language models become the front door to political information, what they refuse to discuss becomes a new instrument of information control. We argue that a model's guardrail encodes not a universal notion of harm but the political threat model of the state that governs its developer, and we derive the expected structure of that control from the comparative study of how authoritarian regimes censor. Across ten models and three languages, Chinese guardrails carry its signatures: they answer to the developer's own regime, refusing identical collective-action prompts far more when a prompt names China than a foreign state; within politics they target the capacity to coordinate rather than dissent, declining even to help organize pro-government mobilization; and their strictness is porous, collapsing under adversarial paraphrase, so that the models most resistant to attack are Western frontier systems, not the strictest refusers. Machine censorship thus reproduces the friction-based logic of prior-era information control while, lacking a censor's case-by-case judgment, proving blunter than the bureaucracy it resembles---so that audits which measure refusal directly overstate how controlled a model actually is.