Cardamom dataset supports detailed Arabic dialect speech recognition

CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

Computation and Language

Summary

Speech recognition systems often struggle with the many subtle dialects spoken across Arabic-speaking countries. The authors created Cardamom, a new collection of about 40 hours of Arabic speech from 21 closely related dialects, each carefully labeled. This dataset helps computers better understand and adapt to these subtle dialect differences, improving transcription accuracy. They tested existing speech systems on Cardamom and showed that adapting them with this data reduces errors significantly.

What this means in practice

  • For speech recognition developers: Improve transcription accuracy by adapting systems to specific micro-dialects within Arabic speech using Cardamom’s fine-grained data.
  • For voice assistant engineers: Build voice assistants that better understand regional Arabic dialects by training on Cardamom’s detailed dialect labels.

Authors

Bashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh, Abdurrahman Juma, Sharaf Makahleh, Nour Gamal, Omar Attia, Hanaa Kurdi, Najwa Rizk, Maysa Anaya, Hessah Altimyat, Layal Alhazmi, Shumukh Alotaibi, Hajar Alhadaris, Rayan Alomari, Rahaf Almalaq, Malak Alkhorasani, Sara alghamdi, Rahaf Alshamrani, Nsrin Ashraf, Ibrahim Jaradat, Nada Qardahji, Yasmin Zaraket, Elmoukhtar Brahim, Sidi Ebeidy, Oumoulmouminin Mahmoud, Yahjeb Bouha Khatraty, Meya Haroune, Mohammad Ghaddar, Mohamad Eldirany, Rashed Alamoush, Tala Chhaytle, Nuha Albadi, Yahya El Hadj, Hamzah Luqman, Fadi A. Zaraket, Mustafa Jarrar, Muhammad Abdul-Mageed

Abstract

We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.