The Effect of Multi-Lingual and Keyword Adversarial Injection on LLM Relevance Judgment
2026-07-11 • Information Retrieval
Information Retrieval
AI summaryⓘ
The authors studied how large language models (LLMs), used to judge the relevance of search results, can be tricked by special types of input called prompt injections in different languages. They tested attacks that add misleading instructions or content in 8 languages and found these attacks can make the models give incorrect high relevance scores while avoiding current defenses. Even improved defense methods can be bypassed by adapting the attacks. This shows there is a big problem with how well these models can be trusted across languages and points to the need for better safeguards.
Large Language ModelsPrompt InjectionInformation RetrievalRelevance EvaluationMultilingual NLPAdversarial AttacksTREC Deep LearningPrompt EngineeringModel RobustnessCross-lingual
Authors
Nguyen Khoi Vo, Duy Duong Tuong, Oleg Zendel, Mark Sanderson
Abstract
Large language models (LLMs) are increasingly being used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation remains insufficiently understood, particularly in multilingual settings. In this work, we investigate the impact of cross-lingual prompt injection attacks on LLM-based relevance judgments using TREC Deep Learning collections and two open-weight models under established prompting frameworks. We examine both instruction-based and content-based injection strategies in 8 languages spanning different resource levels. Our results demonstrate that multilingual query-based injections are highly effective in inflating relevance scores while simultaneously evading existing prompt-injection defenses. We further found that, although existing defense mechanisms can be modified to mitigate such attacks, these injections can be easily adapted to bypass them. These findings highlight a critical gap in current defense approaches and demonstrate that language generalization can act as an attack vector, underscoring the need for more robust and proactive evaluation frameworks for LLM-as-a-judge systems.