Emoji rating bias hides true differences among language models
Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure
Computation and Language
Summary
Measuring how well language models generate emotional emoji summaries in multiple languages is tricky because human ratings vary a lot depending on the reviewer. The authors show that differences between language models mostly disappear when factoring in individual annotator preferences. They also find the number of emojis used affects scores more than quality. Instead of current rating methods, the authors propose a more stable way to measure emoji-based emotion generation.
What this means in practice
- •For multilingual nlp engineers: Improve evaluation of emotion-focused emoji generation systems by using stable decodability metrics instead of inconsistent human preference scores.
- •For ai model evaluators: Design more reliable benchmarks for comparing language model outputs across languages and providers by controlling for annotator bias and output length.
Authors
Fardeen Sadab, Adib Sakhawat
Abstract
We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other ($F(7,14)=0.59$, $p=0.76$), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7\% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by $3.1$ points and change the top-ranked system. In place of preference scoring we propose **emoji-affect decodability**, a reference-based probe whose rankings are stable to $\pm0.003$ macro-F1 across seeds.