Two methods control false alerts when screening AI text inside documents

Two Conformal Constructions for Adaptive Within-Document AI-Text Screening

Computation and Language

Summary

Detecting text generated by AI inside documents can raise false alarms, wasting effort. The authors propose two ways to carefully decide when to stop checking parts of a document to control false alerts, even if tokens inside the document depend on each other. They prove that their methods keep false alarms under control while allowing early stopping, but testing their accuracy and speed in practice is still needed.

What this means in practice

  • For content moderation teams: Ensure AI-generated text detectors maintain false alert rates when checking parts of documents before fully reading them.
  • For digital forensics teams: Use adaptive text inspection with controlled false alert guarantees when screening documents for AI-generated content.

A theory result. No direct application yet.

Authors

Marco Mandap, Jerahmeel Hipolito, Arcel Galvez, Charlie Margaret Balagtas, Michael Joshua Buluran, Jeff Roel Durmiendo, Rizzette E. Lopez

Abstract

We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.