AI tutors match human experts in GRE learning gains at low cost

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Artificial IntelligenceComputers and Society

Summary

Getting better at GRE questions usually needs a good tutor, either a person or a smart computer. The authors created StudentBench, a tool to test if AI tutors can help students learn just as well as human tutors. They found that some AI tutors helped students improve as much as human experts but were much cheaper and faster. This means AI might support learning in new ways without losing quality.

What this means in practice

  • For online education platforms: Integrate AI tutors that provide GRE practice and feedback with the same effectiveness as human tutoring but at significantly reduced costs.$Commercial implications: Allows education platforms to offer scalable AI tutoring services competitively priced compared to human tutors.
  • For education technology developers: Use StudentBench to evaluate and improve AI tutoring systems’ lesson planning and practice problem generation to optimize student engagement and learning.

Authors

Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner, Jonas Mueller

Abstract

Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.