Benchmarks evaluating ai facts often miss global knowledge perspectives
Whose Facts Count? A Culturally Responsive Audit of LLM Evaluation Benchmarks
Human-Computer Interaction
Summary
Many tests used to evaluate large language models (LLMs) rely mostly on English-language facts and sources, which means they don’t fully represent knowledge from around the world. The authors checked popular evaluation sets and found a bias towards certain regions, like Colombia, and English, even though internet users are globally diverse. They created a framework to measure how well these tests account for different cultures and found major gaps. This matters because the tests influence decisions in education, jobs, and public services worldwide.
What this means in practice
- •For ai developers: Improve evaluation tests for language models by including culturally diverse and non-English knowledge to produce fairer assessment results worldwide.
- •For policy makers: Use the cultural responsiveness framework to select or design better benchmarks when making decisions based on AI capabilities affecting diverse populations.
Authors
Fatima Tuz Zahra, Md. Sajeebul Islam Sk., Rachel Chung
Abstract
LLM benchmarks function as evaluation instruments, informing decisions that affect education, labor, and public services worldwide. Drawing on Hood, Kirkhart, and Hopson's culturally responsive evaluation (CRE) frameworks, this paper applies a six-dimension CR rubric to audit OpenAI's SimpleQA (N = 4,326 items) and the LMSYS Chatbot Arena (N = 600 conversations). Every SimpleQA question requires English-language archival verification as its evidentiary basis. A single rater's preoccupation with Colombian founding dates accounts for 2.70% of items, inflating the appearance of Global South coverage. English-language prompts constitute 76.3% of Arena conversations, against an International Telecommunication Union (ITU)-estimated 25.9% share of global internet users. A 50-item counter-benchmark scored a mean CR deficit nearly three times lower than SimpleQA (Cohen's d = 1.01). The paper proposes a practical CR evaluation framework. These are structural validity failures, not incidental measurement problems, with direct consequences for communities whose knowledge traditions these instruments were not built to see.