Papers for

platform content moderators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Gpt models transform gender bias instead of reducing it over versions

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

Abstract: Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($ρ= +0.55$, $p = .034$) while Detoxify does not ($ρ= -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.

Thu 17 SeptComputation and LanguageArtificial Intelligence
The gist
Evaluations that say newer AI language models are less harmful might be misleading. The researchers found that these models often hide or change discriminatory content instead of removing it, a process they call harm laundering. For example, early models showed harmful gender stereotypes clearly, but later models express bias in more subtle ways not caught by standard toxicity checks. This means that using toxicity scores alone is not enough to judge if a model is truly less biased or harmful. The authors suggest a new method to better detect this hidden bias across different AI generations.
Open 2609.20779v1