Image pixels show limits of identifying image origins under attack
Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
Cryptography and SecurityArtificial IntelligenceComputer Vision and Pattern RecognitionMachine Learning
Summary
People want to know if just looking at the pixels in an image can tell where it came from, like if it was made by a person or an AI. The authors studied this as a problem where images can be changed before checking where they came from, making it tricky. They found the best possible accuracy depends on how different the original and edited images are, not the tools used. They also show why current public tools often fail when images are slightly altered and suggest testing the theoretical limits separately from how much information the tool reveals.
What this means in practice
- •For security teams: Assess vulnerability of image verification tools to pixel-level attacks for improved authenticity checks.
- •For image moderation teams: Design better moderation workflows by understanding the limitations of current image origin verification under image editing.
Authors
Kai Yao
Abstract
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.