Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study
2026-08-24 • Computation and Language
Computation and LanguageArtificial Intelligence
AI summaryⓘ
The authors worked on translating between English and Pnar, a language spoken by about 400,000 people that doesn't have many digital resources. They gathered over 10,000 sentence pairs from a local newspaper to train machine translation models. Their best system translated Pnar to English better than the other way around, partly because of differences in sentence structure. They also found that some tuning methods didn't help much due to limited data and identified challenges like new words and mixed language use. This study sets the first baseline for machine translation involving Pnar.
Pnar languagemachine translationparallel corpusBLEU scorestatistical machine translationword order (SOV vs SVO)lexicalized reorderingminimum error rate training (MERT)out of vocabulary (OOV)code mixing
Authors
Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh
Abstract
Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural language processing (NLP) resources. This paper presents the first machine translation study for the English and Pnar language pair. Using articles collected from the Wyrta newspaper, we built a parallel corpus comprising of 10,234 sentences and trained phrase-based statistical machine translation (SMT) systems the models using 9,563 parallel corpora under three configurations for each direction using Moses, GIZA++ , KenLM, varying lexicalized reordering and minimum error rate training (MERT) tuning. The models are evaluated on a held out test set of 371 sentences, the best performing system achieves a BLEU score of 14.97 (chrF2: 33.42, TER: 77.60) for Pnar to English and 11.16 (chrF2: 31.38, TER: 93.51) for English to Pnar, establishing the first quantitative benchmark for this language pair. Lexicalized reordering improves translation quality by 3.73 BLEU points for Pnar to English, reflecting the structural shift from the source language's SOV word order to the target language's SVO order, whereas MERT tuning degrades BLEU performance under low resource conditions. Finally, we analyze the remaining translation errors, including morphological out of vocabulary (OOV) words, long-distance reordering and Khasi code mixing and discuss future directions toward neural and multilingual machine translation for Pnar.