Papers for

localization teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

ChindaMT improves Thai-English translation following detailed instructions

Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation

Abstract: Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.

Mon 28 SeptComputation and Language
The gist
Translating between Thai and English accurately and following specific instructions about word choice and style is hard. The authors created a new method that picks the most useful training examples and makes sure the translated text meets all the given rules from examples that already do. Their resulting translation models, called ChindaMT, produce better or equally good translations compared to existing models, especially when strict rules need to be followed. They also share their trained models, data, and tools for others to use.
Open → 2609.34770v1

Cross-language study reveals how players judge video games differently

Deciphering the Babel of Play: A Human-AI Collaborative Approach for Large-Scale Cross-Language Analysis of Game Reviews

Abstract: We present a large-scale cross-language analysis of game reviews using a human-AI collaborative framework that combines quantitative screening with multilingual large language models (LLMs). Starting from 17 million Steam reviews across 30 languages and 2,000 top-selling titles, we select 28 games with notable cross-language rating patterns. We then apply LLM-assisted content analysis to 442,162 reviews spanning 17 languages, with human researchers guiding codebook development and interpreting the results. Our findings reveal differences in both the aspects language communities prioritize and how they evaluate them, highlighting the roles of narrative expectations, game mechanics and stability, localization quality, cultural proximity, and perceptions of developers and publishers. We also identify rare cases of cross-language consensus. This work offers empirical insights into cross-cultural game evaluation and a scalable methodological approach to multilingual content analysis that preserves human interpretation.

Sat 19 SeptHuman-Computer Interaction
The gist
People from different language backgrounds often rate video games differently because they focus on different things like story, game mechanics, or how well the game is translated. The authors looked at hundreds of thousands of game reviews in many languages using a mix of human help and AI language tools. They found both differences and some rare areas where players agree regardless of language. This helps us understand how culture affects opinions about games.
Open → 2609.23104v1

Gender bias affects machine translation scores across jobs and languages

Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations

Abstract: Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.

Fri 18 SeptComputation and Language
The gist
When translating texts where a person's gender is unknown, machine translation systems and their evaluation tools sometimes prefer masculine forms over feminine ones. The authors studied how this bias appears in translations involving different occupations across several languages. They found that masculine translations often get higher scores, and these preferences sometimes reflect job gender stereotypes. However, this bias varies widely depending on the language and the evaluation method used.
Open → 2609.21490v1

TeochewBench benchmark evaluates Teochew Hanzi translation quality

TeochewBench: A Human-Reviewed Benchmark for Teochew Hanzi Translation

Abstract: Teochew has a substantial speaker community and exhibits distinctive lexical, syntactic, and pragmatic features, yet textual resources for evaluating large language models remain limited. We present TeochewBench, a human-reviewed benchmark comprising 300 Teochew Hanzi expressions for evaluating translation from Teochew Hanzi into Mandarin Chinese and English. The dataset covers five categories: basic vocabulary; everyday sentences; Teochew-specific expressions; tone, politeness, and context; and idiomatic, ambiguous, and culturally specific expressions. A primary Teochew-speaking reviewer examined all entries individually and revised them as needed, while two additional Teochew speakers verified selected items. Our main evaluation covers 11 official general-purpose post-trained models on the reviewed dataset in both translation directions, yielding 6,600 predictions. Two official base checkpoints provide 1,200 predictions for supplementary diagnostics, bringing the total to 13 models and 7,800 predictions. We additionally include a Hanzi-copy control, which returns the source input unchanged, to assess how shared Hanzi affect automatic scores for translation into Mandarin Chinese. Qwen3.5-27B achieved the highest overall chrF-style score among the evaluated checkpoints, at 60.63, followed by Qwen2.5-72B-Instruct at 56.61, Gemma-3-27B-IT at 56.36, and GLM-4-32B-0414 at 55.82. Across the 11 main-evaluation models, the mean chrF-style score decreased from 69.25 for low-specificity items to 27.52 for high-specificity items. High-specificity expressions received lower scores and exhibited smaller cross-model differences, suggesting that they constitute a shared low-scoring region across the model families evaluated here. The Hanzi-copy control further indicates that surface overlap in low-specificity items can substantially affect automatic scores for translation into Mandarin Chinese.

Wed 16 SeptComputation and Language
The gist
Teochew is a distinct Chinese dialect with unique words and expressions, but there are few resources to test how well computers translate its written form. The authors created TeochewBench, a collection of 300 Teochew written phrases checked by native speakers, covering everyday and culture-specific expressions. They tested 13 language models to see how accurately they translate these phrases into Mandarin and English. Results showed models struggle more with specific or culturally unique phrases, highlighting translation challenges where meanings are subtle or tied to context.
Open → 2609.18156v1

Measuring translation quality loss from mixing similar Mozambican language varieties

Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+

Abstract: In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja with Chichewa, and evaluate NLLB-200, Google Translate, GPT, and a variant-aware NLLB model. Holding system output fixed reveals substantial reference sensitivity. On \textit{devtest}, changing only the reference from Tsonga to Xichangana reduces spBLEU by 13.10 points for NLLB-200 and 15.30 for Google. On matched Nyanja subsets, replacing Chichewa with Mozambican Nyanja produces smaller but consistent reductions of 3.03 and 6.10 spBLEU, respectively. Variant-aware fine-tuning reverses this pattern on the intended targets: relative to NLLB-200, it improves Xichangana by 7.04 spBLEU and Mozambican Nyanja by 5.33 on \textit{devtest}, while losing performance on the sibling references. GPT is competitive on Tsonga and Chichewa but substantially weaker on the Mozambican varieties. For Sena, the finetuned model reaches 12.64 spBLEU and 36.21 chrF++ on \textit{devtest}. These findings motivate variety-aware language identifiers, references, and reporting for cross-border languages or language dialects/variants. The data is publicly available on Hugging Face at https://huggingface.co/datasets/MOZNLP/FLORES_MOZ

Sat 12 SeptComputation and Language
The gist
When automatic translation tools try to translate languages with many similar varieties or dialects, mixing them up can cause big drops in quality. The authors added new test data sets for three Mozambican languages to an existing benchmark to study this problem. They found that using the wrong variety as a reference for evaluation can reduce measured translation accuracy by up to 15 points. Training models to recognize specific language varieties improved results for the intended targets but harmed accuracy on related varieties. This work highlights the need for language-aware tools and more precise evaluation when dealing with closely related languages.
Open → 2609.13847v1

Multilingual large language models struggle with Urdu stories

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.

Wed 9 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Large language models (LLMs) can write stories in many languages, but their ability in less common languages like Urdu is not well understood. The authors studied three popular LLMs that generated Urdu stories and found many problems. These stories often had basic grammar mistakes, confusing meanings, repeated phrases, and lacked cultural depth. Even when given extra examples to improve, the issues remained. This shows that current LLMs are not yet reliable for creating or understanding content in low-resource languages like Urdu.
Open → 2609.10758v1

Machine translation accuracy improves by accounting for term variation

Improving Term Evaluation in Machine Translation: Variation Matters

Abstract: Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term. However, human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to account for this variation in document-level MT evaluation of English-French scientific translation, combining glossary-based accuracy, translation consistency, and a new cross-term variation (CTV) diagnostic measure that tests whether variation relationships are preserved across languages. Based on analyses of two parallel corpora, translated by four MT systems, we find that (1) MT systems generate less target-side variation than human translators; (2) transfer patterns strongly depend on the variation type; (3) consistency rankings vary with the choice of metric; and (4) constraining MT with a glossary improves accuracy and consistency but degrades CTV by suppressing valid variation. We argue for variation-aware evaluation that conditions consistency penalties on whether target-side variation mirrors source-side variation.

Tue 8 SeptComputation and Language
The gist
Machine translation systems usually treat each word or term as having only one correct translation, but human translators often use different valid versions of the same term. The authors studied English-to-French scientific translations and found that current evaluation methods wrongly penalize translations for these natural variations. They introduced a new way to measure whether variations in the translation reflect variations in the source text. Their work shows that recognizing valid variations can make evaluation of translation quality more accurate and fair.
Open → 2609.08779v1