Arabic language models get faster dynamic tests for facts
AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic
Computation and Language
Summary
Many big language models can speak and understand Arabic better than before, but it's hard to know how much true information they really have, especially about Arab culture and history. The authors made a new tool called AraDynFact that quickly checks what these models know by automatically making questions from Arabic Wikipedia facts. They used this tool to test several models and found it agrees well with older, slower tests. This means AraDynFact can help quickly and cheaply measure how much these models really know about Arabic topics.
What this means in practice
- •For arabic nlp developers: Automatically generate and run fact-checking tests on Arabic language models to validate their accuracy without heavy manual work.
- •For content moderation teams: Ensure Arabic language AI tools provide culturally accurate information by quickly auditing their factual knowledge coverage.
Authors
Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi
Abstract
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains largely underexplored. Current evaluation metrics often focus on translation or generic reasoning, failing to capture the rich historical, social, and regional nuances inherent to Arabic culture. In addition, most benchmarks rely on heavy work, with human intervention in some steps, making the evaluation of knowledge coverage expensive and slow. To address this deficiency, we introduce AraDynFact, a novel dynamic evaluation framework designed to rigorously assess the factual Arabic knowledge embedded in LLMs. Unlike static benchmarks, AraDynFact employs a dynamic approach to extract factual information and generate rich and answerable questions in a fast and automatic way. We apply AraDynFact to Arabic Wikipedia and audit the performance of several state-of-the-art models, ranging from Arabic-centric specialized LLMs to high-resource general purpose LLMs. In addition we found a high degree of correlation with existing, hand-crafted Arabic-centric benchmarks, confirming the potential of our dynamic approach.