ChindaMT improves Thai-English translation following detailed instructions
Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation
Computation and Language
Summary
Translating between Thai and English accurately and following specific instructions about word choice and style is hard. The authors created a new method that picks the most useful training examples and makes sure the translated text meets all the given rules from examples that already do. Their resulting translation models, called ChindaMT, produce better or equally good translations compared to existing models, especially when strict rules need to be followed. They also share their trained models, data, and tools for others to use.
What this means in practice
- •For localization teams: Produce Thai-English translations that consistently follow customer style and terminology guidelines with high accuracy.$Commercial implications: Enables translation service providers to offer more precise, instruction-compliant outputs improving client satisfaction in language localization.
- •For software developers: Integrate fine-tuned Thai-English models to improve bilingual app interfaces and documentation translations with controlled output style.
Authors
Thodsaporn Chay-intr, Krittapad Harnchang, Mahannop, Thabua, Kobkrit Viriyayudhakorn, Thanaruk Theeramunkong
Abstract
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.