DeepEdu improves speed and accuracy of Vietnamese AI tutoring systems

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

Artificial Intelligence

Summary

AI tutoring could help students learn better in places like Vietnam, but common approaches have problems with privacy and accuracy on local subjects. The authors created DeepEdu-v1, which works faster and better on Vietnamese curriculum by using a new way to handle long learning conversations and by learning from past interactions without retraining. This reduces data privacy concerns and makes the AI more accurate on complex tasks. Their system cuts response times and improves correctness compared to existing methods.

What this means in practice

  • For education technology developers: Build AI tutoring platforms that respect local data laws and handle large textbook-based conversations efficiently.
  • For chatbot platform engineers: Improve speed and accuracy of chatbots on region-specific content by integrating self-updating verified knowledge without costly model retraining.

Authors

Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu

Abstract

AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.