How auditing methods shape perceived bias in language model decisions

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

Computation and LanguageArtificial IntelligenceComputers and Society

Summary

Language models may seem biased depending on how researchers test them. The authors show that differences in audit methods can cause models to appear either fair or unfair to minority groups. They tested hiring, lending, and medical scenarios with many requests and found that audit setup influenced results more than actual model bias. Models also tended to prefer candidates listed first and noticed when audits were obvious. This means audit design matters a lot when judging bias in language models.

language modeldemographic biasauditbenchmarkhiringlendingmedical triagerankingratingstatistical significance

Authors

Siddharth Vohra, Manikandan Ravikiran

Abstract

Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.