Papers for

machine translation developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Representation guides faster and better parallel text generation

Representation-based Masked Diffusion Model

Abstract: Masked Diffusion Models (MDMs) have emerged as a compelling paradigm for language modeling, offering the capability for efficient parallel text generation. However, existing parallel sampling methods typically update multiple masked tokens independently and ignore the complex mutual dependencies among the masked tokens. This independent updating mechanism lacks global coordination and might lead to incoherent outputs. To address this limitation, we propose Representation-based Masked Diffusion Model (RMDM), a framework that leverages the text representation to explicitly encode global semantics and help to parallel update tokens more precisely. Specifically, we first encode text into a continuous semantic space using a pretrained encoder and learn an invertible transformation that normalizes the representation distribution to a Gaussian prior, facilitating efficient sampling during generation. Conditioned on this latent semantic representation, we train a masked diffusion model to learn the conditional text distribution, where the representation serves as global semantic guidance to coordinate parallel token updates and faithfully approximate the target distribution. Empirical results demonstrate that RMDM significantly improves generation quality, particularly in aggressive few-step sampling regimes.

Fri 11 SeptComputation and Language
The gist
Generating sentences where many words change at once can confuse computers because those words depend on each other. The authors found a way to help computers understand the overall meaning of the sentence first and then update the words together more smoothly. They do this by turning the sentence into a math-friendly code that keeps track of the big picture. With this, their method creates clearer and more accurate sentences, especially when changing many words quickly.
Open 2609.12382v1

Large language models produce noisy translations cleaned by new benchmark

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevalence, this problem lacks dedicated benchmarks and systematic study. We analyze over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identify 12 recurring noise patterns, which we group into formatting and content noise. Building on the observed patterns, we construct TransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs, comprising 8,800 synthetically generated instances and 1,100 manually curated authentic instances. We evaluate two extraction approaches on the TransClean benchmark: 1) a span-based extraction method leveraging translation quality estimation models for span detection, and 2) an LLM-based extraction method that prompts an LLM to isolate the translation. Our benchmark and analysis provide the first systematic framework to evaluate and improve the cleanliness of LLM translation outputs.

Thu 10 SeptComputation and Language
The gist
Sometimes when large language models translate text, they add extra words like labels or repeated phrases that aren’t part of the actual translation. The authors studied many examples of these noisy translations and found common patterns of this extra text. They created TransClean, a dataset to help test methods that can automatically remove the unwanted parts and keep only the clean translation. They also compared two ways to extract the clean translations: one that detects important parts using quality estimates and another that asks a language model to find the right text.
Open 2609.11399v1

Arabic morphological generation remains difficult for large language models

YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

Abstract: Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.

Wed 9 SeptComputation and Language
The gist
Arabic words can change their form in many ways, which is tricky for language technology. The authors created YallaMorph, a big test set that helps check how well computer models can produce these word forms properly. They found that even advanced models struggle with some Arabic word forms, especially rare or combined ones. This shows there is still work to do for better Arabic language tools.
Open 2609.10153v1

EviSI improves evaluation of real time speech translation quality

EviSI: An Evaluation Agent for Simultaneous Interpreting

Abstract: Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the source stream continues. To support timely delivery and limit accumulated delay, systems adopt reformulation and summarization, which can preserve meaning while departing from written references. BLEU and COMET may not reliably distinguish such variation from semantic loss. We introduce EviSI, a large language model evaluation agent adapting the error analysis and penalty principles of Multidimensional Quality Metrics (MQM). It constructs shared source evidence, assesses semantic fidelity and oral expression, reconciles overlapping errors and scores deterministically. EviSI recovers the aggregate human system ranking for English to Chinese. Mean Kendall agreement with human system rankings within corpora reaches 0.707 for English to Chinese and 0.467 for Chinese to English, exceeding evaluated baselines. An extension across five directions shows positive concordance with COMET without human ratings. Individual output agreement with humans remains mixed.

Tue 8 SeptComputation and Language
The gist
Real-time speech translation has to understand, translate, and speak while someone is still talking. Usual ways to check if the translation is good may get confused because spoken translation can change words to keep up without losing meaning. The paper’s authors created EviSI, a tool that uses a language model to better judge if these translations keep the meaning and sound right. EviSI matches human opinions about translation quality better than some current tools, especially for English to Chinese.
Open 2609.08171v1

Korean scholarly abstracts show rising AI style changes after 2023

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

Abstract: Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018-August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1-2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: sisahada "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like araboda "look into" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024-2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025-2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent.

Mon 7 SeptComputation and LanguageDigital Libraries
The gist
This paper studies changes in Korean academic writing from 2018 to 2026 by looking at words that appear more often than expected after 2022. The authors found a rise starting in late 2024 in words that suggest a style shift linked to large language models (LLMs), like AI assistants. These changes are partly visible even after accounting for topic differences and aren’t explained by translations. Similar patterns appear earlier in English abstracts of the same articles. The study helps show how AI tools might be influencing how academic papers are written in Korean.
Open 2609.07447v1

Language models struggle more with rare words but keep some grammar skills

FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

Abstract: Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones. However, these benchmarks largely ignore lexical frequency variation, despite lexical frequency being a pervasive and highly skewed property of natural language use. Consequently, existing evaluations do not test whether grammatical preferences remain stable when contrasts involve rare lexical items. We introduce FreqBLiMP, a frequency-controlled extension of BLiMP that regenerates all 67 paradigms under explicit Zipf-frequency regimes while preserving each minimal-pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, we find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood, but only a modest reduction in overall contrastive acceptability accuracy. However, this aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena that require lemma-specific information.

Mon 7 SeptComputation and LanguageArtificial Intelligence
The gist
Language models are tested to see if they can tell good sentences from nearly identical bad ones. The authors found that when sentences have rare words, these models give them lower probabilities but still mostly pick the right grammar differences. However, the models do worse with grammar that depends on specific word knowledge and stay strong only on common grammar rules. This new test helps reveal where language models are sturdy and where they become fragile with unusual vocabulary.
Open 2609.07153v1

Line coupled language model speeds up token generation per step

Line-Coupled Language Model

Abstract: Autoregressive language models generate one token per decoding step, limiting the useful output of each forward pass. Although diffusion models, insertion-based decoding, and multi-token prediction enable parallel generation, they either incur additional training-time token traffic or struggle to predict strongly dependent future tokens. We introduce the Line-Coupled Language Model (LCLM), an autoregressive model that advances multiple text lines together by predicting the next token for every active line while coupling the lines through shared causal context. LCLM interleaves line tokens into a single causal sequence and uses line-staggered rotary positions, retaining the standard next-token objective and causal attention. Controlled experiments show that cross-line targets are substantially less dependent than consecutive same-line targets, supporting lines as parallel generation units. With 881M parameters, LCLM produces an average of 2.94 content tokens per forward pass with a validation cross-entropy loss of 2.44, compared with 1.00 token per forward pass and a loss of 2.39 for the vanilla autoregressive baseline. Most notably, even when LCLM generates 16 tokens per forward pass, its loss is only 0.09 higher than that of the vanilla autoregressive baseline (2.34 vs. 2.25).

Mon 7 SeptComputation and Language
The gist
Generating text one word at a time makes language models slow because each step only creates one word. The authors introduce a new model called the Line-Coupled Language Model (LCLM) that predicts multiple next words across different lines at the same time, while still using information from all lines together. This lets the model produce almost three times as many words per step with only a small drop in accuracy. Their approach keeps the usual way of training and understanding text sequences but speeds up the process by cleverly arranging how lines of text are predicted together.
Open 2609.07129v1

Dynamic programming improves byte pair encoding vocabulary selection

Dynamic-Programming-Guided Hierarchical BPE and Empirical Analysis of Vocabulary Pruning

Abstract: Byte Pair Encoding (BPE) constructs vocabularies through greedy pair merging, but the resulting merge order does not necessarily allocate a fixed model-visible vocabulary optimally for compression. We propose Dynamic-Programming-Guided Hierarchical BPE (DH-BPE), a vocabulary-construction method that combines token exposure under exact minimum-token segmentation with the hierarchical dependencies induced by BPE training. Starting from a modestly overshot BPE candidate vocabulary, DH-BPE uses dynamic programming to measure candidate utility and applies exposure-guided, dependency-aware pruning to select a fixed-size model-visible vocabulary. We compare DH-BPE against Standard BPE and recent vocabulary-optimization baselines, including Pruned BPE, MinGram, and MinGram-PP, in primary evaluations at 12K and 16K target vocabulary sizes, with an additional 18K evaluation against MinGram only. Across the primary 12K and 16K comparisons, DH-BPE consistently improves aggregate compression over Standard BPE, Pruned BPE, and MinGram under a shared exact minimum-token DP encoder. MinGram-PP achieves stronger aggregate compression in the primary comparisons, but DH-BPE outperforms it at overshoot factors f = 2.0 and f = 3.0 in cross-corpus evaluation; at 12K, MinGram-PP reverses this ordering only with the substantially larger candidate pools at f = 4.0 and f = 5.0. Qualitative analysis further shows that DH-BPE balances later, more complete BPE merges with reusable subword components, providing a practical approach to improving vocabulary allocation under a fixed model-visible vocabulary budget.

Mon 7 SeptComputation and LanguageMachine Learning
The gist
When computers turn text into pieces to understand it better, they often use a method called Byte Pair Encoding (BPE). The order in which BPE chooses these pieces isn't always the best for compressing text efficiently. The authors came up with a new method, DH-BPE, that uses a smart planning technique called dynamic programming to pick a better set of pieces. Their tests show this method can compress text more efficiently than several existing methods, especially when carefully selecting how many pieces to consider at first.
Open 2609.06898v1