Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit

2026-07-20Computation and Language

Computation and LanguageArtificial Intelligence
AI summary

The authors looked at ways to check the quality of test responses in multiple languages without having to translate them into English first. They tested different computer models that turn sentences into numbers (called embeddings) to see if these could do the same job as translations in judging reliability. Their results showed that using these native-language embeddings gave similar reliability results and even recovered some responses lost due to translation problems. Overall, this means it might be possible to assess multilingual responses directly without relying on translations.

multilingual assessmentsentence embeddingsLinguistic-Integrated Reliability Auditing (LiRA)PIRLSconstructed-response itemstranslationreliability estimationembedding models
Authors
Ummugul Bezirhan, Ji Yoon Jung, Matthias von Davier
Abstract
Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability Auditing (LiRA) across 11 PIRLS constructed-response items and three embedding models. Native-language embeddings reproduced translation-based reliability estimates closely while recovering responses excluded after translation failure, with no meaningful change in reliability.