Vision language models improve handwritten exam grading accuracy and reliability

Vision-Language Models for Criterion-Level Grading of Handwritten Examinations in Outcome-Based Education

Computer Vision and Pattern Recognition

Summary

Grading handwritten exams by hand takes a lot of time and can differ between markers. The authors study if vision-language AI models can grade exams by matching performance to specific learning goals. They tested different AI setups on nearly 2,000 exam answers and found some models agree with human graders better than humans agree with each other. However, the AI grading varied between runs and explanations for grades were not clearly helpful. This shows that AI can assist in grading but needs careful tuning and review by teachers.

What this means in practice

  • For academic assessment teams: Deploy vision-language AI to reduce manual grading workload while maintaining alignment with learning outcomes in handwritten exam scoring.
  • For education technology developers: Incorporate calibrated AI grading modules into digital assessment platforms to assist teachers with reliable, criterion-level scoring of scanned handwritten exams.$Commercial implications: The paper enables development of AI grading products that improve scoring efficiency and consistency in educational tools.

Authors

Md Khalid Syfullah, Asif Hasan Tonmoy, Saad Ahmed, S. M. Jahangir Alam

Abstract

Criterion-level grading connects examination performance to learning outcomes, but manual marking introduces workload and variation between markers. This study evaluates vision-language models (VLMs) for handwritten outcome-based assessment across five dimensions: accuracy, human agreement, repeated-run reliability, error concentration, and explanation quality. Using 1,982 criterion-level records from 485 undergraduate examination answers, we compare 20 configurations spanning Qwen2.5-VL, InternVL3, Pixtral, a Donut baseline, and a cascade ensemble. Evaluation setups include zero-shot prompting, few-shot prompting, partial fine-tuning, and Low-Rank Adaptation (LoRA). Two independent faculty markers regraded all 291 test criteria, providing a human agreement baseline on the same assessment materials. Qwen2.5-VL with LoRA achieved Quadratic Weighted Kappa (QWK) of 0.727 and mean absolute error of 0.435 marks against the examiner, compared with mean human-pair QWK of 0.551. This comparison reflects calibration to the examiner's training marks. LoRA outperformed partial fine-tuning for all three instruction-tuned VLMs, while few-shot prompting reduced QWK in every configuration with valid prompted scores. Aggregate reliability and exact repeatability diverged: intraclass correlations ranged from 0.790 to 0.874, yet 50.2-63.6% of criteria changed marks across five sampled runs. Attention-guided deletion showed no statistically significant advantage over random masking, and four faculty reviewers reached no consensus on explanation usefulness. These findings highlight the need for rubric-specific calibration, repeatable scoring, review of consequential errors, and separate validation of explanations. The released evaluation protocol supports criterion-level assessment research and grading tools with teacher oversight.