Set-level attacks can mislead retrieval-augmented generation systems

LENS: The Sum Is Worse Than the Parts for Set-Level Poisoning in Retrieval-Augmented Generation

Cryptography and Security

Summary

Retrieval-augmented generation (RAG) systems use many documents together to answer questions. The authors found a way to create tricky sets of documents that seem harmless by themselves but together mislead the system into giving wrong answers. They built a method called LENS that carefully creates these tricky sets and tested it on several systems, showing it works better than earlier methods. This reveals a new way RAG systems can be attacked and helps test how well defenses work.

What this means in practice

  • For ai security teams: Evaluate and improve defenses for retrieval-augmented generation systems against subtle set-level poisoning attacks.
  • For software developers: Test the robustness of document-based AI systems by generating complex adversarial document sets using a multi-agent framework.

Authors

Kaisheng Fan, Yishu Gao, Xunzhu Tang, Tegawend'e F. Bissyand'e, Weizhe Zhang

Abstract

Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper subset remains insufficient in frozen single-round RAG. We formalize set-level compositional poisoning, where documents designed to remain individually plausible jointly redirect RAG outputs to a target answer, while proper subsets fail to induce the target on their own. To construct such attacks, we propose LENS, a generator-black-box multi-agent framework that casts construction as constrained evidence composition. LENS factorizes target inference into a query-conditioned interpretation lens and complementary facts, then uses a nested dual-loop workflow to concentrate steering in the full set while suppressing subset leakage. The outer loop plans the interpretation lens and semantic roles; the inner loop synthesizes documents and applies counterexample-guided repair. Across four benchmarks and three generators, returned packets achieve 0.852 full-set ASR and 0.784 post-retrieval ASR@5, while their strongest proper subsets reach only 0.069. Against construction baselines evaluated on the same frozen manifest, LENS improves all-attempt E2E-Strict@5 from 0.244 to 0.363, a 48.8% relative gain. A blinded human audit finds that 68.3% of returned packets combine an incorrect target, a definite answer-criterion shift, and no target entailment under the original semantics. Across four published defenses, LENS attains the highest defended all-attempt ASR@5, exceeding the strongest baseline by 0.141 on average. Together, these results establish evidence composition as a distinct RAG security boundary and position LENS as a stress test for defenses that reason over document sets.