LLM based agents get better defense against hidden prompt attacks

CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

Cryptography and SecurityArtificial Intelligence

Summary

Large language models (LLMs) used in smart agents can get tricked by hidden instructions that make them do harmful things without clear warnings. The authors introduce CoDeL, a method that trains the agent and a set of attackers together in a cycle so the agent learns from the latest tricky attacks. This approach helps the agent spot and refuse hidden harmful prompts even when they are disguised over multiple steps. Their tests show CoDeL reduces how often attacks succeed by a large margin compared to other methods.

What this means in practice

  • For ai system developers: Improve safety of LLM-based assistants by training them against evolving hidden prompt attacks that mimic realistic workflows.
  • For security teams in tech companies: Evaluate and update defenses for AI agents exposed to complex prompt injection threats by using CoDeL’s co-evolutionary attacker-defender training.

Authors

Xiao Yang, Yangchen Ou, Yuhan Gao, Le Wang, Zonghao Ying, Aishan Liu

Abstract

Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender is updated each round via LoRA-based GDPO under a decoupled reward over safety, task progress, and format compliance, so refusing injections and completing the user's task jointly define fitness. To keep supplying it with the failures worth learning from, a co-evolving prober searches over injection rounds, attack methods, and payloads for injections that still penetrate the current defender, guided jointly by attack success and attack latency so that it preferentially mines breaches the defender notices too late. Each defender update invalidates part of the attack population and forces the next round onto a new frontier, turning the defender's own failures into a moving curriculum. Extensive experiments on three IPI benchmarks, nine baselines, and two base models show that CoDeL reduces attack success rate (ASR) by 88.5% and outperforms other baselines largely (+38.0%). Codes are available.