Arabic speech large language models gain new training resources
Nuha-Speech: Building General-Purpose Arabic Speech-LLMs
Computation and Language
Summary
Arabic speech technology is behind other languages because there are fewer tools and datasets available. To help fix this, the authors created Nuha-Speech, which includes a large set of Arabic spoken questions and answers for teaching computers to understand speech. They used this dataset to improve existing speech language models with training focused on Arabic. They also built a way to test these models across different speech tasks to see how well they work. This work aims to build the basic tools needed for better Arabic speech AI despite limited existing resources.
Speech large language modelArabic speech recognitionDataset constructionInstruction tuningSupervised fine-tuningQuestion answeringEvaluation frameworkMultilingual models
Authors
Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi
Abstract
As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap, we introduce Nuha-Speech, a comprehensive initiative to develop general-purpose Arabic speech-LLMs spanning dataset construction, model training, and systematic evaluation. Specifically, we constructed a large-scale Arabic Speech Question-Answering (SQA) corpus comprising over 1.5 million training samples to allow instruction tuning over a broad range of core speech tasks. Then, the corpus was used for supervised fine-tuning based on Qwen-Omni model variants at different scales. Finally, we designed an evaluation framework featuring diverse tasks and tailored metrics. Through this work, we aim to establish foundational infrastructures for Arabic Speech-LLMs under constraints imposed by limited Arabic speech resources.