ChatGPT images dataset reveals challenges in detecting image generator versions

ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Online images made with AI tools like ChatGPT Images 2.5 can keep the same name even when the generator behind them changes, making it hard to tell which version created each image. The authors collected over 3,400 images shortly after ChatGPT Images 2.5 launched and studied where the images came from and what they showed. They tested six different tools designed to detect AI-generated images and found big differences in how well these tools worked, showing that detection is still tricky. This dataset helps understand how image creation tools evolve and how to better detect AI-generated content.

What this means in practice

  • For content moderation teams: Identify and flag AI-generated images during rapid updates of underlying generators using source and attribution clues in a launch period dataset.
  • For digital forensics analysts: Evaluate the reliability of AI image detectors under changing model versions for verifying image origin authenticity.

Authors

Dennis Ng, Xingyu Shen, Ankit Raj, Kidus Zewde, Tommy Duong, Yuchen Zhou, Yuxin Zhang, Neo Tiangratanakul, Simiao Ren

Abstract

An image tool can change its underlying generator while retaining its public name, making version attribution from online posts ambiguous. We study this problem after the ChatGPT Images 2.5 launch. Our frozen collection contains 3,478 images from 2,440 posts across 8 sources. Recorded posting times fall within the first 51.1 hours after the announcement. It records three attribution tiers and retains standalone images after image-form filtering and targeted review. Caption claims and host records provide admission evidence, not independently verified generator identity. The observed content profile depends on the source mixture: NightCafe supplies 39.0% of images but 77.0% of CLIP-assigned fantasy scenes. We then evaluate six frozen detectors at thresholds calibrated to a 5% flag rate on reference photographs. Collection flag rates range from 3.7 to 56.4%, falling 42-81 percentage points below GenImage recall. Held-out artwork false-positive rates range from 1.5 to 96.5%, so a higher collection flag rate does not by itself establish better detection. An exploratory X-only comparison with our April collection finds a higher September flag rate for Effort, and a suggestive difference for DoU, under fixed-threshold post-clustered bootstrap intervals. Attribution, content and processing differences prevent a causal interpretation of these contrasts. The collection supports analysis of reported model use during a product transition, with source and attribution evidence retained for interpretation.