Gender bias affects machine translation scores across jobs and languages
Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations
Computation and Language
Summary
When translating texts where a person's gender is unknown, machine translation systems and their evaluation tools sometimes prefer masculine forms over feminine ones. The authors studied how this bias appears in translations involving different occupations across several languages. They found that masculine translations often get higher scores, and these preferences sometimes reflect job gender stereotypes. However, this bias varies widely depending on the language and the evaluation method used.
What this means in practice
- •For machine translation developers: Improve MT evaluation tools by detecting and mitigating gender bias in scoring translations across occupations and languages.
- •For localization teams: Use findings to better interpret MT output quality scores where gender representation may bias translation assessments in different languages.
Authors
Orfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed, Giuseppe Attanasio, Chrysoula Zerva
Abstract
Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.