PSMP-CLIP improves zero-shot anomaly detection with better image masks

PSMP-CLIP: Patch-Prompt SAM and Multi-Semantic Prompting for CLIP-Based Zero-Shot Anomaly Detection

Computer Vision and Pattern Recognition

Summary

Detecting unusual parts in images without using examples from that image type is hard. Existing methods using a model called CLIP struggle to pinpoint odd spots precisely. The authors present PSMP-CLIP, which helps by creating better image masks from smaller image patches and by using multiple guided text prompts for improved detection. They tested this approach on many datasets and found it works better than similar methods in identifying anomalies at the pixel level.

What this means in practice

Authors

Xuezhi Xiang, Guanghao Wu, Heqi Xiang, Jiayao Liu, Xiaoheng Li, Yiming Chen, Shanjun Zhang

Abstract

Zero-shot anomaly detection aims to localize anomalies without target-domain samples. Existing CLIP-based methods suffer from coarse anomaly maps and limited semantic prompts. We propose PSMP-CLIP, integrating patch-prompt SAM2 segmentation (PPSS) and multi-semantic guided prompt regularization (MSGPR). PPSS samples prompts directly from intermediate patch features, avoiding threshold drift and guiding SAM2 to produce precise masks. MSGPR uses multiple learnable prompts constrained by semantic anchors to preserve generalization. Experiments on 14 datasets show highly competitive performance, achieving the best pixel-level AUROC on MVTec AD, BTAD, DTD-Synthetic, CVC-ClinicDB, TN3K, Endo, and Kvasir.