Arabic morphological generation remains difficult for large language models

YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

Computation and Language

Summary

Arabic words can change their form in many ways, which is tricky for language technology. The authors created YallaMorph, a big test set that helps check how well computer models can produce these word forms properly. They found that even advanced models struggle with some Arabic word forms, especially rare or combined ones. This shows there is still work to do for better Arabic language tools.

What this means in practice

Authors

Mahmoud Reda, Salam Khalifa, Reham Marzouk, Nizar Habash

Abstract

Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.