Does ChatGPT score research quality differently by gender?

2026-08-10Digital Libraries

Digital Libraries
AI summary

The authors studied whether ChatGPT, an AI language model, rates research papers differently based on the first author's gender using nearly 90,000 UK research articles. Although author gender was hidden from ChatGPT, papers with male first authors scored slightly higher in most scientific fields. This male bias was stronger in ChatGPT scores than in the official research rankings, but differences were small and not seen in social sciences or humanities. The authors suggest these differences might come from other factors like subject area or journal, not writing style or direct AI bias. They caution that AI-based research evaluation should be used carefully.

Large Language ModelsChatGPTResearch EvaluationGender BiasUK Research Excellence FrameworkUnits of AssessmentAuthorshipAI BiasAbstract ComplexityDepartmental Averaging
Authors
Kayvan Kousha, Mike Thelwall
Abstract
Large Language Models (LLMs) are being considered for research evaluation, raising concerns about the introduction of AI bias. This study investigates whether ChatGPT research quality scores differ by first-author gender using 89,744 journal articles from the UK Research Excellence Framework (REF) 2021. Author information was withheld from ChatGPT to avoid direct gender bias. Nevertheless, male first-authored papers had slightly higher ChatGPT scores in most Units of Assessment (UoAs), especially in health, science and engineering-related subjects, and this pattern was often stronger for ChatGPT than for REF scores, based on a departmental-level proxy. Rank-based ChatGPT gains relative to REF scores were also more favourable for male first-authored papers in most UoAs, although the differences were generally small. Gender differences were not evident for solo research in the social sciences, arts and humanities, however. The male-favouring pattern for first-authored research was not explained by gender differences in writing styles, at least as reflected in abstract complexity. Some ChatGPT-REF differences may also reflect the departmental averaging process used to generate the REF proxy scores. Average ChatGPT scores may differ by first-author gender indirectly through other factors, such as field, topic, method, journal context or authorship structure. Thus, this is an additional reason to be cautious with AI-based research evaluation.