LLMs detect AI-generated text with varying success across versions

Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis

Computation and LanguageArtificial Intelligence

Summary

Distinguishing text written by people from text made by AI language models is important but tricky. The authors tested 15 different AI models of various ages to see how well they can spot AI-written text, including text from themselves and from other models. They found that newer AI detectors do a better job overall, but none consistently detect their own text better than others. Also, different model versions make different types of errors when trying to decide if a text is human or AI written. Finally, the way AI models explain their decisions varies a lot.

What this means in practice

  • For content moderation teams: Identify AI-generated content across different AI model versions to improve detection accuracy and consistency in online platforms.
  • For digital forensics analysts: Use insights on generational detection biases to assess text authenticity and better understand AI-generated text artifacts.

Authors

Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang, Ruizhe Li, Shujun Li

Abstract

Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self-detection versus cross-detection across model generations, remains poorly understood. We systematically evaluate 15 LLMs spanning three model generations as both generators and detectors. Using a benchmark of 1,000 human-written texts and 15,000 LGTs (1,000 per model), we collected over 233,000 binary classifications alongside natural-language explanations. Our results reveal that detection efficacy is primarily driven by detector capability rather than generator provenance, although outputs from newer generators remain notably harder to detect. Crucially, statistical comparisons show no systematic advantage or disadvantage for self-detection across models. Error analysis further exposes generational bias shifts: first-generation detectors under-detect LGTs (high false-negative rates), second-generation detectors over-flag human texts (high false-positive rates), and the latest models achieve balanced trade-offs. Finally, we highlight significant inconsistencies in how different LLMs apply textual cues to justify their decisions. Code: https://github.com/hyyuan/detect-llm-generated-texts.