Measuring translation quality loss from mixing similar Mozambican language varieties

Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+

Computation and Language

Summary

When automatic translation tools try to translate languages with many similar varieties or dialects, mixing them up can cause big drops in quality. The authors added new test data sets for three Mozambican languages to an existing benchmark to study this problem. They found that using the wrong variety as a reference for evaluation can reduce measured translation accuracy by up to 15 points. Training models to recognize specific language varieties improved results for the intended targets but harmed accuracy on related varieties. This work highlights the need for language-aware tools and more precise evaluation when dealing with closely related languages.

What this means in practice

  • For machine translation developers: Improve translation accuracy by training and evaluating models on specific language varieties instead of conflated dialects.
  • For localization teams: Create better translation evaluation sets tailored to individual dialects for cross-border or related languages to ensure quality.

Authors

Felermino D. M. A. Ali, Delfina Lázaro Mateus, Manuel Valente Mangue

Abstract

In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja with Chichewa, and evaluate NLLB-200, Google Translate, GPT, and a variant-aware NLLB model. Holding system output fixed reveals substantial reference sensitivity. On \textit{devtest}, changing only the reference from Tsonga to Xichangana reduces spBLEU by 13.10 points for NLLB-200 and 15.30 for Google. On matched Nyanja subsets, replacing Chichewa with Mozambican Nyanja produces smaller but consistent reductions of 3.03 and 6.10 spBLEU, respectively. Variant-aware fine-tuning reverses this pattern on the intended targets: relative to NLLB-200, it improves Xichangana by 7.04 spBLEU and Mozambican Nyanja by 5.33 on \textit{devtest}, while losing performance on the sibling references. GPT is competitive on Tsonga and Chichewa but substantially weaker on the Mozambican varieties. For Sena, the finetuned model reaches 12.64 spBLEU and 36.21 chrF++ on \textit{devtest}. These findings motivate variety-aware language identifiers, references, and reporting for cross-border languages or language dialects/variants. The data is publicly available on Hugging Face at https://huggingface.co/datasets/MOZNLP/FLORES_MOZ