Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

2026-08-10Computation and Language

Computation and LanguageArtificial Intelligence
AI summary

The authors explain that most multilingual translation tests start from English and translate into other languages, which can cause problems and miss local cultural differences. They suggest a new way to test translations called source-contrastive evaluation, and show this using Cultivar, a localized version of the FLORES dataset. By comparing localized and non-localized translations, they can spot issues like data contamination and how well models handle different cultures. After testing 32 translation models, they found that models specialized in machine translation were less reliable, some models seemed to overfit the benchmark, and most models translated US English better than other locales.

multilingual translationbenchmarksource-contrastive evaluationlocalizationdata contaminationFLORES datasetmachine translation modelsoverfittinglocalecultural considerations
Authors
Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
Abstract
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.