AI summaryⓘ
The authors studied how much help large language models (LLMs) give when students use them as tutors during real learning tasks. They created a five-level scale to measure how directly LLM responses assist students, from minimal hints to fully solving problems. Analyzing over 14,000 responses in a university AI course, they found most answers gave high levels of help, either explaining or solving tasks. While the level of help influenced how students continued conversations, it didn't strongly predict their later exam scores beyond what was predicted by past performance and dialogue behavior. The authors provide a way to measure and compare the kind of assistance LLM tutors offer in educational settings.
Large Language ModelsTutoringScaffoldingEducational AIStudent PerformanceConversational BehaviorNatural Language ProcessingAssessmentHuman-Annotated DataDialogue Systems
Authors
Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin
Abstract
Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.