Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?
2026-08-17 • Computation and Language
Computation and Language
AI summaryⓘ
The authors studied how well advanced language models help people learn English by checking if they can find and fix common mistakes and explain why they're wrong. They tested different model settings and also used extra information to improve answers. While the models did a good job correcting sentences, their explanations weren't detailed or clear enough for real teaching. This shows that even though using AI in language classes sounds promising, we still need to better understand how good these tools really are at teaching.
Large Language ModelsCorrective FeedbackLanguage PedagogyRetrieval-Augmented GenerationGLEUBERTScoreLinguistic NuanceCultural SensitivityInstructional AppropriatenessEnglish Language Learning
Authors
Kristina Šekrst, Ana Kovačić
Abstract
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.