Papers for

multilingual nlp engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Emoji rating bias hides true differences among language models

Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure

Abstract: We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other ($F(7,14)=0.59$, $p=0.76$), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7\% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by $3.1$ points and change the top-ranked system. In place of preference scoring we propose **emoji-affect decodability**, a reference-based probe whose rankings are stable to $\pm0.003$ macro-F1 across seeds.

Thu 24 SeptComputation and Language
The gist
Measuring how well language models generate emotional emoji summaries in multiple languages is tricky because human ratings vary a lot depending on the reviewer. The authors show that differences between language models mostly disappear when factoring in individual annotator preferences. They also find the number of emojis used affects scores more than quality. Instead of current rating methods, the authors propose a more stable way to measure emoji-based emotion generation.
Open → 2609.29445v1

Structural similarity improves cross language scientific sentence classification

Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

Abstract: Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.

Thu 17 SeptComputation and LanguageDigital Libraries
The gist
Scientific papers often organize sentences into sections like introduction, methods, and results, which follow similar patterns across languages. This paper shows that using the structural similarities of these patterns helps computers better understand and classify sentences in scientific papers written in different languages. The authors created a multilingual dataset and found that language closeness matters less than how sentences are organized. They developed methods that use this structure information and improved classification accuracy when transferring models to languages without training data.
Open → 2609.19650v1