Public Kurdish speech files have accuracy and labeling problems

Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study

Computation and Language

Summary

Many people speak Kurdish, but technology to read it aloud is limited. The authors checked recently released Kurdish speech recordings and found problems like mislabeled test files and incorrect equipment records. These defects don’t change reported results but could make it harder for others to use the files to build better Kurdish voices. Kurdish voices in the files reflect only a few speakers’ reading styles, not the full variety of Kurdish speech.

What this means in practice

  • For software developers: Use the public Kurdish voice recordings to build text-to-speech systems that consider labeling and data quality issues identified in the analysis.
  • For community linguists: Review and correct the public speech data to improve Kurdish voice quality and enable broader reuse in regional dialect modeling.

A survey. It maps existing work.

Authors

Hiwa Asadpour

Abstract

Kurdish is spoken by millions of people, but little technology can read it aloud. A recent study released three Kurdish voices, 35 hours of recorded speech, and a paper describing the work, all free to download. This review checks how well those public files match the paper. The research is careful about its limits, but the files contain several problems: a settings file lists equipment that was never used, test recordings are left unlabeled among training data, and a coding fault mishandles long numbers. The download page also claims a stronger result than the paper reports and recommends one voice for general use. That recommendation matters because Kurdish has major regional and written variation, while these voices were built from three people reading prepared texts. The process therefore removes much everyday and regional speech. English and German benefit from long traditions of dictionaries and linguistic description that help identify wrong pronunciations; Kurdish has far less such support, so software choices can go unchecked. The voices sound fluent, but they represent the reading styles of their speakers rather than Kurdish as a whole. Most of these issues can be fixed using information the team already has, without changing the reported results. Better records would mainly make the work easier for others, especially community linguists, to check and reuse. The license is the main exception: whether audiobook owners allow corrected versions to be shared will affect whether future Kurdish voices can build on this work or must start again.