Arabic dialect bias causes worse language model performance on sa'idi
Some Dialects Are More Equal Than Others: Non-Prestigious Arabic Dialectal Bias in LLMs
Computation and Language
Summary
Many AI language models understand Egyptian Arabic mainly through its common urban dialect, Cairene Egyptian Arabic, while the rural Sa'idi dialect is rarely represented. The authors found that this lack of exposure causes these models to prefer Cairene features over Sa'idi ones, even when Sa'idi usage is correct. This bias lowers the model's performance on tasks involving Sa'idi Arabic, showing that not all dialects are treated equally by current language technology. The study points to the need to better include and support dialects like Sa'idi in language AI.
What this means in practice
- •For natural language engineers: Improve language models by including less-represented dialects like Sa'idi to reduce bias and performance drops on diverse Arabic language inputs.
- •For multilingual ai developers: Develop more inclusive multilingual systems by correcting dialectal biases that affect understanding and generation in underrepresented Arabic sub-dialects.
Authors
Mai Mohamed Eida, Ryan Dolan, Paul de Nijs, Jonathan Dunn
Abstract
Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representation for the less prestigious Sa'idi Egyptian Arabic (SEA) dialect both in LLM and resource development. Does this lack of representation influence an LLM's view of the acceptability of SEA (upstream), and does an upstream bias against SEA lead to worse performance (downstream)? We investigate the upstream effect of SEA dialectal features on LLM preferences in a Targeted Syntactic Evaluation (TSE) task which reveals a significant bias against SEA across multiple LLMs. We then analyze the effect of these same features on downstream model performance on MMLU benchmarks and show that models experience a degradation in performance when presented with SEA. This work highlights the need for further exploration on how sub-dialectal variation impacts language technologies.