Machine translation evaluation improves by considering word choice variation

Improving Term Evaluation in Machine Translation: Variation Matters

Computation and Language

Summary

Machine translation systems are often judged by how well they translate specific terms, but usually only one correct translation is expected. The authors show that human translators naturally use different valid versions of translations, which current evaluation methods unfairly mark as mistakes. They propose a new approach that checks if the variation in translations matches the variation found in the original text, providing a more accurate way to evaluate translation quality. Their study with English to French scientific texts shows that forcing machine translation to stick to one term improves accuracy but reduces natural variation. Overall, they suggest evaluation should consider matching variation rather than penalizing it blindly.

machine translationterm evaluationtranslation consistencyglossarycross-term variationscientific translationEnglish-French translationparallel corpora

Authors

Nicolas Dahan, Ziqian Peng, François Yvon, Rachel Bawden

Abstract

Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term. However, human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to account for this variation in document-level MT evaluation of English-French scientific translation, combining glossary-based accuracy, translation consistency, and a new cross-term variation (CTV) diagnostic measure that tests whether variation relationships are preserved across languages. Based on analyses of two parallel corpora, translated by four MT systems, we find that (1) MT systems generate less target-side variation than human translators; (2) transfer patterns strongly depend on the variation type; (3) consistency rankings vary with the choice of metric; and (4) constraining MT with a glossary improves accuracy and consistency but degrades CTV by suppressing valid variation. We argue for variation-aware evaluation that conditions consistency penalties on whether target-side variation mirrors source-side variation.