QuranicMMLU benchmark tests AI skills on Quranic Arabic language
QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge
Computation and Language
Summary
Evaluating AI on the Quran’s Arabic language is tricky because it involves complex linguistic features. The authors created QuranicMMLU, a new test with questions covering phonology, morphology, syntax, semantics, and pragmatics of Quranic Arabic. They organized questions by cognitive difficulty and verse complexity, and checked AI answers both automatically and manually. Testing 12 AI systems showed that special Islamic models do best, and multiple-choice tests often overestimate AI understanding compared to open-ended answers. This benchmark helps measure Quranic Arabic language skills more accurately in AI.
What this means in practice
- •For arabic nlp developers: Evaluate and improve Arabic natural language processing models specifically on Quranic linguistic knowledge using a detailed, cognitively-aware benchmark.
- •For language technology product teams: Develop Quran-focused AI tools like educational apps or digital assistants with more reliable assessments of language understanding.
Authors
Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari
Abstract
We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguistic competencies or stratifying by cognitive demand and verse difficulty. We construct a five-pillar Quranic taxonomy spanning Phonology, Morphology, Syntax, Semantics, and Pragmatics, with 31 leaves covering phenomena from tajwīd and root-and-pattern morphology to occasions of revelation and inter-surah coherence. For each leaf we generate questions stratified by Bloom's cognitive level and verse perplexity, then have LLM as a judge to independently answer and score every item and route the annotations to manual review. The resulting dataset comprises 980 human-reviewed questions, each issued in both open-ended and multiple-choice form. We benchmark 12 systems on these items and find that the Islamic-specialized model leads, yet every system scores higher on multiple-choice accuracy (average 84%) than open-ended answer quality (average 60%): the two rankings agree closely (Kendall's τ=0.73), but multiple-choice scoring hides failures that surface only once answer choices are removed. QuranicMMLU thus offers a rigorous, linguistically grounded framework for evaluating Arabic NLP in the Quranic domain.