Email agents detect prompt injection attacks by modeling attack stages
Prompt Injection Detection for Email Agents Through Attack Chain Modeling
Cryptography and SecurityComputation and Language
Summary
Large language model email assistants can be tricked by harmful instructions hidden inside emails. The authors developed a method that watches for these tricks by looking at the attack as a series of steps, not just spotting bad words. Their method combines text checks, rules, and user behavior analysis to catch suspicious activity. They found that training the system with tricky but safe emails helps reduce false alarms while still detecting real attacks.
What this means in practice
- •For email security teams: Improve email assistant defenses by detecting multi-stage prompt injection attacks, reducing the chance malicious emails influence automated responses.
- •For chatbot developers: Design AI assistants that better analyze user requests and contextual email content for suspicious commands, lowering risks of indirect prompt injection.
Authors
Ahmad Hashmi, Dhyey Patel, Yunting Yin
Abstract
Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy. To support this framework, we derive attack chain labels from prompt injection datasets, evaluate the proposed framework under random splits, temporal phase transfer, conditional stage transfer, cross-dataset transfer, and conduct ablation studies on multiple benchmarks. Results show that random train test splits substantially overestimate robustness under distribution shift, while later tool argument stages are more predictable than earlier stages in the framework. We also show that training on harmless emails that resemble attacks helps reduce false alarms while preserving the ability to detect real attacks. Across five binary benchmarks, our framework achieves a mean F1 score of 0.406 under the strict threshold setting policy, compared with 0.216 for the strongest of five pretrained detectors evaluated without additional training. These results highlight the value of combining attack stage predictions with checks for conflicts between the user's request and instructions in retrieved emails. Our experiments also demonstrate the importance of training with challenging benign examples to balance attack detection and false alarms.