Framework improves math tutoring by learning from past conversations

REAT: A Reflective Experience-Augmented Tutoring Framework for Multi-turn Mathematical Instruction

Multiagent Systems

Summary

Solving math problems well doesn’t always mean a computer can teach math effectively. The authors designed a system called REAT that learns from previous tutoring chats and uses what it has learned to help students better during live sessions. It watches past conversations to find good teaching patterns and then uses those at the right moments, adapting to how each student is doing. Tests show that this method helps more in tough teaching moments than just giving the computer fixed instructions or retraining it.

What this means in practice

Authors

Jianheng Zhou, Chaoli Zhang, Xingjun Wei, Xinliang Zhou, Giancarlo Fortino, Xing Fan, Yanfeng Wang, Qingsong Wen, Haoyang Li

Abstract

Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.