Papers for

image moderation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

CLIPGuard defends image AI from hidden embedding space backdoors

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

Abstract: Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git

Fri 25 SeptComputer Vision and Pattern RecognitionCryptography and Security
The gist
Some bad actors hide secret triggers in images that fool AI models like CLIP, making them behave wrongly. Existing defenses often need to see inside the model or lots of clean data, which is not always possible. The authors designed CLIPGuard, a tool that treats the model as a black box and finds suspicious parts of an image to fix them without hurting the good parts. Their tests show CLIPGuard can stop these hidden backdoors effectively while keeping the AI’s normal accuracy high.
Open → 2609.31558v1

Image pixels show limits of identifying image origins under attack

Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance

Abstract: Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.

Fri 25 SeptCryptography and SecurityArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
People want to know if just looking at the pixels in an image can tell where it came from, like if it was made by a person or an AI. The authors studied this as a problem where images can be changed before checking where they came from, making it tricky. They found the best possible accuracy depends on how different the original and edited images are, not the tools used. They also show why current public tools often fail when images are slightly altered and suggest testing the theoretical limits separately from how much information the tool reveals.
Open → 2609.30997v1