Visual difference guided few-shot anomaly detection improves results
VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
Artificial Intelligence
Summary
Finding unusual objects or defects in images often means carefully comparing a new image to normal ones. The authors note that simply describing differences using language misses many small visual details. They introduce a method called VD-DeepStack that merges detailed visual difference information with language reasoning to spot anomalies more accurately, especially when only a few examples are available. Their experiments on industrial and medical datasets show that this approach outperforms previous methods relying mainly on text-based comparisons.
What this means in practice
- •For quality control engineers: Detect subtle defects in manufacturing by comparing new product images with few normal references more accurately.
- •For medical imaging technicians: Identify rare anomalies in medical scans with limited healthy examples by improving fine-grained visual inspection combined with language reasoning.
Authors
Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang, Yifang Xu, Yuchuan Wu, Xiaolei Chen, Zhengtao Yao, Fan Shi, Yang Liu, Bin Li, Xiangyang Xue
Abstract
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.