Large language models show bias when assessing students with different backgrounds

The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment

Computation and LanguageArtificial IntelligenceComputers and Society

Summary

Large language models (LLMs) are used to grade and give feedback to students. This paper finds that these models notice and react to information about a student's background, like education level, whether that information is given directly or implied in conversation. Sometimes, this helps by making feedback easier to read for students with less education. But it can also cause unpredictable unfair biases, like giving more negative responses to students with lower education. The authors tested six top LLMs on tasks like essay scoring and question answering to show how these biases appear.

What this means in practice

Authors

Donya Rooein, Luca Benedetto, Dirk Hovy

Abstract

Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. However, it also risks being a cause of discrimination, e.g., when assigning lower scores to students from lower socioeconomic backgrounds. We set up controlled prompts to test 1) explicit demographic effects, where we mention demographic details directly, and 2) implicit effects, where we use conversation history as a demographic signal. We test these settings in three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. We test six state-of-the-art LLMs on these tasks. In both explicit and implicit cases, the models pick up on demographic cues and can change their scoring, feedback, and answers accordingly. We find that LLMs frequently adjust the readability of feedback to education levels when these are explicitly mentioned. On the other hand, implicit conditions produce unpredictable biases, such as in question answering, where responses from lower-education levels receive lower sentiment scores. Our results provide clear evidence of demographic sensitivity in LLMs for educational assessment tasks.