Large language models show bias when assessing students with different backgrounds
The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment
Computation and LanguageArtificial IntelligenceComputers and Society
Summary
Large language models (LLMs) are used to grade and give feedback to students. This paper finds that these models notice and react to information about a student's background, like education level, whether that information is given directly or implied in conversation. Sometimes, this helps by making feedback easier to read for students with less education. But it can also cause unpredictable unfair biases, like giving more negative responses to students with lower education. The authors tested six top LLMs on tasks like essay scoring and question answering to show how these biases appear.
What this means in practice
- •For education technology developers: Adjust LLM-based student feedback tools to better handle explicit demographic data without introducing unfair bias.
- •For customer support teams: Monitor conversational AI for unintended demographic biases when implicit user background signals influence responses.
Authors
Donya Rooein, Luca Benedetto, Dirk Hovy
Abstract
Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. However, it also risks being a cause of discrimination, e.g., when assigning lower scores to students from lower socioeconomic backgrounds. We set up controlled prompts to test 1) explicit demographic effects, where we mention demographic details directly, and 2) implicit effects, where we use conversation history as a demographic signal. We test these settings in three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. We test six state-of-the-art LLMs on these tasks. In both explicit and implicit cases, the models pick up on demographic cues and can change their scoring, feedback, and answers accordingly. We find that LLMs frequently adjust the readability of feedback to education levels when these are explicitly mentioned. On the other hand, implicit conditions produce unpredictable biases, such as in question answering, where responses from lower-education levels receive lower sentiment scores. Our results provide clear evidence of demographic sensitivity in LLMs for educational assessment tasks.