Multilingual large language models struggle with Urdu stories

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Computation and LanguageArtificial IntelligenceMachine Learning

Summary

Large language models (LLMs) can write stories in many languages, but their ability in less common languages like Urdu is not well understood. The authors studied three popular LLMs that generated Urdu stories and found many problems. These stories often had basic grammar mistakes, confusing meanings, repeated phrases, and lacked cultural depth. Even when given extra examples to improve, the issues remained. This shows that current LLMs are not yet reliable for creating or understanding content in low-resource languages like Urdu.

What this means in practice

  • For content moderators: Detect and flag linguistic and cultural errors in automatically generated Urdu text to maintain quality control.
  • For localization teams: Assess limitations of popular LLMs when generating content in Urdu to improve translation or adaptation strategies.

Authors

Farah Adeeba, Abdul Rafae Khan, Rajesh Bhatt, Hassan Sajjad

Abstract

Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.