De-identified résumés still reveal ethnicity beyond language details
Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits
Computation and LanguageArtificial Intelligence
Summary
Removing obvious personal information like language skills from résumés does not fully stop AI models from guessing a person’s ethnic background. The authors tested multiple AI models on résumés where they controlled language details to be the same but changed other writing aspects. Even without language clues, the AI could often identify ethnic groups, especially when subtle hints in the text were strong. They also found that how evaluation tests are designed can greatly change the results, showing that measuring bias needs careful methods.
What this means in practice
- •For human resources teams: Improve fairness audits by recognizing that ethnic information can leak from résumé text beyond explicit fields like language skills.
- •For ai model evaluators: Design bias evaluation protocols that account for cue salience and evaluation artifacts to better assess demographic inference risks.
Authors
Qiangju Chen, Yang Xiao
Abstract
De-identified résumé screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual résumés. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\ge94\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from résumé content from effects introduced by the evaluation protocol.