Papers for

email security teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Phishing emails use varied themes but mainly urge clicking links

A Large-Scale Empirical Study of Modern Phishing Email Content

Abstract: Phishing remains one of the most pervasive threats to Internet users, and email remains its predominant delivery channel. Email content is the attack surface of phishing: it is what the victim reads and what automated defenses inspect. Yet the composition of modern phishing content is poorly measured. Prior work has characterized dimensions such as theme, call-to-action (CTA), and impersonation, but not at scale, and their associations and temporal changes remain unclear, owing to small or source-specific corpora, bag-of-words topic models, and a focus on text alone. We present a content-focused measurement study of 2.9M distinct real-world phishing emails collected over 13 months (June 2025 - June 2026) in collaboration with the Anti-Phishing Working Group (APWG). We treat each email as a composite artifact comprising message text and its attachments: 272K images, 143K PDFs, and 57K calendar invitations. Using an LLM pipeline validated against human-annotated samples, we analyze these components along three dimensions (theme, CTA, and impersonation), examine the associations among them, and measure longer-term change against a historical dataset. We find that attackers diversify what they use to deceive but converge on how victims should respond: no theme exceeds 21.3% of emails, while a single CTA, URL navigation, accounts for 73.0%. CTA and impersonation choices are conditioned on theme. Attachments play three roles: images supplement the message text, PDFs substitute for it by carrying the pretext, and calendar invitations reinforce it by replicating interaction endpoints into a persistent medium. Over the longer term, the dominant CTA for invoice-themed phishing shifted from URL navigation to offline communication, rising from 6.7% in 2015 to 46.9% in 2025.

Fri 25 SeptCryptography and Security
The gist
Phishing emails trick people by pretending to be trustworthy sources, and most come through email. This study by the authors looked closely at nearly 3 million real phishing emails to understand what these messages say and how they try to fool victims. They found that while phishing emails cover many topics, most of them ask people to click on a link. The study also shows how attackers use images, PDFs, and calendar invites in different ways to support their tricks, and how phishing tactics have shifted over time.
Open → 2609.30683v1

Email agents detect prompt injection attacks by modeling attack stages

Prompt Injection Detection for Email Agents Through Attack Chain Modeling

Abstract: Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy. To support this framework, we derive attack chain labels from prompt injection datasets, evaluate the proposed framework under random splits, temporal phase transfer, conditional stage transfer, cross-dataset transfer, and conduct ablation studies on multiple benchmarks. Results show that random train test splits substantially overestimate robustness under distribution shift, while later tool argument stages are more predictable than earlier stages in the framework. We also show that training on harmless emails that resemble attacks helps reduce false alarms while preserving the ability to detect real attacks. Across five binary benchmarks, our framework achieves a mean F1 score of 0.406 under the strict threshold setting policy, compared with 0.216 for the strongest of five pretrained detectors evaluated without additional training. These results highlight the value of combining attack stage predictions with checks for conflicts between the user's request and instructions in retrieved emails. Our experiments also demonstrate the importance of training with challenging benign examples to balance attack detection and false alarms.

Fri 25 SeptCryptography and SecurityComputation and Language
The gist
Large language model email assistants can be tricked by harmful instructions hidden inside emails. The authors developed a method that watches for these tricks by looking at the attack as a series of steps, not just spotting bad words. Their method combines text checks, rules, and user behavior analysis to catch suspicious activity. They found that training the system with tricky but safe emails helps reduce false alarms while still detecting real attacks.
Open → 2609.30657v1

Clipping improves detecting AI-generated text even with editing errors

Robust Detection of LLM-Generated Text under Contamination

Abstract: We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.

Thu 24 SeptMachine Learning
The gist
Detecting whether text is written by a human or an AI can be tricky when the AI-generated text is altered or mixed with other content. The authors studied this problem by modeling both human and machine text with a mathematical tool and found the limits where detection is possible. They show that by modifying existing detection tests with a simple 'clipping' step, the tests become more reliable against tricky edits or contamination. Their experiments on several datasets and models confirm that clipping consistently improves the ability to spot AI-generated text while keeping false alarms low.
Open → 2609.29935v1

Gmail extension detects Nigerian fintech phishing with sender checks and AI

A Gmail-Based Phishing Detection Prototype for Nigerian Fintech Emails Using Sender Checks and BiLSTM Classification

Abstract: Phishing emails that impersonate Nigerian fintech providers can combine deceptive sender addresses, lookalike links, and locally familiar language. This study presents a Gmail browser extension that integrates sender-domain and URL checks with a bidirectional long short-term memory (BiLSTM) classifier. The extension compares visible sender addresses and links with profiles for eight fintech platforms, obtains a phishing probability from a locally hosted Flask service, and displays a legitimate, warning, or phishing verdict when an email is opened. The BiLSTM classifier was evaluated on 8,943 test messages from a cleaned dataset of 59,622 phishing and legitimate emails. The test confusion matrix recorded 4,308 true negatives, no false positives, one false negative, and 4,634 true positives. These counts correspond to 99.99% accuracy, 100.00% precision, 99.98% recall, and 99.99% F1 score. Tokenized sequence analysis identified 5.79% overlap between the training and test sets, which may inflate performance estimates for independent messages. A Gmail demonstration showed the integrated extension producing user-visible verdicts, although the complete system was not evaluated on a labeled test set. The findings establish the feasibility of the implemented prototype while leaving its end-to-end detection performance and generalization to unseen attacks open for further evaluation.

Wed 23 SeptCryptography and Security
The gist
Phishing emails trick people by pretending to be from trusted Nigerian fintech companies. The authors made a Gmail extension that checks sender addresses and links, plus uses a special AI model called BiLSTM to spot scam emails. Their tests showed very high accuracy on a large collection of emails, though some data overlap may make the results look better than they are. They demonstrated the extension works in Gmail, but haven’t fully tested it on new unknown phishing emails yet.
Open → 2609.28305v1

Adaptive system improves email spam and phishing detection accuracy

AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection

Abstract: Email spam and phishing attacks remain a critical security threat. Adversaries increasingly exploit large language models to craft contextually convincing malicious messages, and existing spam detection systems often struggle to keep pace. Generalization across diverse and evolving attack scenarios is limited, which reduces effectiveness once these systems are deployed in practice. This paper introduces Adaptive Uncertainty-Routed Analysis (AURA), a multimodal email threat detection system that analyzes both the content of an email and its embedded URLs. AURA is built around two layers: the first quantifies prediction uncertainty from a URL classifier, and only ambiguous messages are escalated to a fine-tuned transformer encoder for semantic analysis. The system is evaluated on eight heterogeneous training corpora together with two held-out real-world corpora spanning a decade of adversarial campaigns. AURA reaches a macro F1-score of 0.9858 in-distribution, and on NazPhish-Eval and GuenterTrap-Eval it maintains 0.9502 and 0.9436, respectively, which is evidence of robust generalization under genuine distribution shift.

Thu 17 SeptMachine Learning
The gist
Spam and phishing emails are dangerous because they try to trick people and can be hard to spot. The authors created a new system called AURA that looks closely at both the message and any links inside it. AURA first checks if the link is suspicious and only digs deeper into the email if it’s unsure. This two-step method helps the system better spot tricky emails, even ones it hasn't seen before, making email safer for everyone.
Open → 2609.19873v1

Subspace interaction improves spotting odd words in text efficiently

SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection

Abstract: Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-grained document-level detection by localizing anomalous tokens within text. By providing fine-grained abnormality prediction, token-level text anomaly detection plays a critical role in various real-world applications, such as spam filtering and fake news detection. However, existing methods still rely on the global distance calculation for scoring, during which the local anomaly signals are severely diluted by numerous redundant normal feature dimensions. Moreover, pre-trained language models used in these methods inevitably smooth out surface anomalies, further limiting their effectiveness in token-level anomaly detection. To address these limitations, we propose a Subspace Interaction-based Method (SIM for short) for token-level text anomaly detection. To prevent local signal dilution, SIM adopts a subspace interaction-based anomaly detector, which decouples high-dimensional token embeddings into multiple low-dimensional ones, amplifying localized anomaly signals hidden within specific dimensions. To counteract the over-smoothing effect, we design a hard pseudo-anomaly generation module to construct pseudo-anomalous tokens, simulating the subtle anomalies obscured by semantic smoothing. Also, a probabilistic boundary loss is developed to standardize anomaly scores into statistical distances, effectively enforcing anomalous instances to deviate significantly from the normal distribution center. Extensive experiments on multiple benchmark datasets verify the effectiveness of SIM and demonstrate its remarkable efficiency, robustness, and interpretability. The source code is available at: https://github.com/yankehan/SIM-TAD.

Tue 8 SeptMachine Learning
The gist
Detecting unusual or suspicious words inside texts is important for spotting things like spam or fake news. Existing methods often miss these odd words because they get lost when looking at the whole text all at once. The authors propose SIM, a new approach that looks at smaller parts of the word representations to better catch strange signals. They also create fake anomalies to help the system learn what unusual words look like despite smooth language models hiding them. Testing shows SIM works well, is fast, and easy to interpret.
Open → 2609.08200v1