Large language models show limited skill in mimicking learner feedback

How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?

Computation and Language

Summary

Understanding if artificial intelligence can mimic how students evaluate educational feedback is important for creating better learning tools. The authors studied how well large language models can simulate real high-school biology students’ feedback judgments. They found that these models struggle to match real learner opinions. Including personal details about learners helps with predicting individual opinions but does not consistently improve agreement on group-level feedback. This suggests more research is needed to find out which types of learner information help AI better capture human preferences.

What this means in practice

Authors

Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara, Kyosuke Takami

Abstract

While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.