Clipping improves detecting AI-generated text even with editing errors
Robust Detection of LLM-Generated Text under Contamination
Machine Learning
Summary
Detecting whether text is written by a human or an AI can be tricky when the AI-generated text is altered or mixed with other content. The authors studied this problem by modeling both human and machine text with a mathematical tool and found the limits where detection is possible. They show that by modifying existing detection tests with a simple 'clipping' step, the tests become more reliable against tricky edits or contamination. Their experiments on several datasets and models confirm that clipping consistently improves the ability to spot AI-generated text while keeping false alarms low.
What this means in practice
- •For content moderators: Improve flagging of AI-generated text even when it is edited or mixed with human text to maintain content quality and authenticity.
- •For email security teams: Enhance detection of AI-generated phishing or spam emails that have been modified to avoid simple detection.
Authors
Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari
Abstract
We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.