The Role of Prompt Language and Translation-Theory-Driven Prompts in Large Language Models: A Case Study on Spanish-Chinese Journalistic Translation

2026-07-03Computation and Language

Computation and LanguageArtificial Intelligence
AI summary

The authors studied how different ways of asking GPT-5.2 to translate Spanish news articles into Chinese affect translation quality. They tested various prompt styles and languages and found that automated scores favored a basic prompt, while human judges preferred a brief-oriented prompt. Using prompts based on translation theory helped reduce certain awkward phrasing errors, but some unnatural language issues remained. The language used in the prompt didn’t significantly change results. The authors suggest these findings could help improve teaching translation, but more user studies are needed.

GPT-5.2prompt engineeringSpanish-Chinese translationBLEU scoreBERTScoreMultidimensional Quality Metrics (MQM)translation theorymachine translation evaluationeditorialsjournalistic translation
Authors
Haohong Lai, Weijia Li
Abstract
This study examines how prompt language and translation theory-driven prompt design influence the quality of Spanish-Chinese journalistic translations generated by GPT-5.2. A parallel corpus of four editorials from El Pais was translated under 48 experimental conditions (4 prompt types, 3 prompt languages, and 4 articles). Translation quality was assessed using BLEU and BERTScore-F1 for automated evaluation, alongside human evaluation based on the Multidimensional Quality Metrics (MQM) framework. Automated metrics identified the baseline prompt (BASE) as the best-performing condition, whereas human evaluation ranked the brief-oriented prompt (BRIEF) highest (MQM: 8.66 vs. 7.84), a reversal likely attributable to the single-reference constraint inherent in automated measures. Sub-error type analysis revealed that translation theory-driven prompts selectively reduced Awkward style errors, while Unidiomatic style errors persisted across conditions. Prompt language had a negligible impact under both evaluation paradigms. These results indicate that translation theory-driven prompts can yield measurable quality gains under expert evaluation of journalistic translations, although their pedagogical implications for language learners remain suggestive and require validation through user-based studies.