Agents improve alignment of classical texts and translations with fewer defects

When Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their Translations

Computation and Language

Summary

Matching ancient texts with their translations helps computers translate and study these works, but it is hard to do accurately. The authors compare several computer methods to line up passages in original texts with their translations. They find that AI agents can recover more matching parts than standard embedding techniques and make fewer mistakes. However, the advantage of agents is smaller when long texts are split into chunks. Having a second check by another AI adds little benefit. Overall, agents produce more reliable and better-aligned outputs.

What this means in practice

  • For machine translation engineers: Improve alignment of classical texts and translations by integrating AI agents to increase accuracy and reduce structural errors in training data.
  • For digital humanities teams: Use AI agents to reliably structure classical texts aligned with their translations, easing computational analysis of ancient manuscripts.

Authors

Máté Metzger

Abstract

Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Tibetan, comprising 9,833 human-aligned units: four embedding pipelines, a direct LLM call, an autonomous agent, and the agent revised by an independent auditor. Generative workflows recover 93-94% of reference correspondences, against at most 77% for embeddings. A ceiling analysis shows that sentence boundaries make some references unrepresentable by the embedding pipelines. Reference recovery is similar across generative workflows: the agent's advantage is 0.5 percentage points (95% CI -0.02 to 1.17), and auditing adds no established benefit. Agents nevertheless produce structurally valid output for all 452 texts, against 437 for direct calls. A blinded three-LLM panel assesses every generative mismatch against the source and human reference. Most mismatches are labelled defensible editorial variation; consensus major-error labels cover only 0.06-0.14% of units. The panel labels significantly fewer residual defects for agents than direct calls (0.7% versus 1.4%), suggesting that reference recovery alone understates alignment quality. On ten long Pali discourses taken as published online, agents and audited agents raise recovery from the direct call's 71% to 84% and 92%. Identical reference-located chunks bring all three to 93%. Agents thus improve structural reliability and reduce judged defects on short passages, while their large recovery advantage on long documents disappears after chunking. In this setting, independent auditing offers little measurable additional benefit on prepared passages.