Papers for

quality control engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Visual difference guided few-shot anomaly detection improves results

VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection

Abstract: Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.

Mon 28 SeptArtificial Intelligence
The gist
Finding unusual objects or defects in images often means carefully comparing a new image to normal ones. The authors note that simply describing differences using language misses many small visual details. They introduce a method called VD-DeepStack that merges detailed visual difference information with language reasoning to spot anomalies more accurately, especially when only a few examples are available. Their experiments on industrial and medical datasets show that this approach outperforms previous methods relying mainly on text-based comparisons.
Open → 2609.34949v1

Automated interpretability improves visual anomaly detector accuracy

Augmenting Visual Anomaly Detection with Automated Interpretability

Abstract: Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different components of this signal. We decompose PatchCore nearest-normal residuals into sparse features using Sparse Autoencoders (SAEs), and provide high-activation and contrastive non-active examples to a Multimodal LLM, which describes each feature and labels it as anomaly, distractor, or uncertain. These labels guide interventions in the SAE hidden representation, where distractor features are suppressed and anomaly features amplified. The edited representation is then used to reconstruct patch embeddings, which are rescored with PatchCore. Across 40 categories from four benchmarks, applying both interventions jointly improves macro-average image-level AUROC from 0.8724 to 0.8857 on source data and from 0.8066 to 0.8210 under synthetic corruptions. On three additional RobustAD categories with real acquisition shifts, the same interventions improve AUROC from 0.8745 to 0.9056 on source data and from 0.6069 to 0.6599 under real acquisition shifts. Finally, individual feature interventions across all 43 categories show that the MLLM labels are aligned in aggregate with how features differently affect normal and anomalous images.

Sun 27 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Visual anomaly detectors find things that look different from normal, but sometimes they confuse unusual but harmless details for real problems. The authors show how to break down the detector’s signals into smaller parts and use a language model to label each part as a real anomaly, a distractor, or uncertain. By amplifying real anomaly signals and suppressing distractors, the system better detects actual problems. This method improves accuracy across many categories and continues to work well even when images are corrupted or changed.
Open → 2609.33818v1

3D inspection system improves fault detection with precise location and type

AT3D-AD: Anomaly Type-Aware 3D Anomaly Detection via Hierarchical Point-Language Alignment

Abstract: Detecting and localizing 3D point-cloud defects is essential for industrial inspection. However, existing methods often suffer from imprecise localization due to the lack of anomaly supervision and reliance on single-granularity representations. To address these limitations, we propose Anomaly Type-Aware 3D Anomaly Detection (AT3D-AD), a unified framework for joint detection, localization, and classification. Specifically, we first design the Physics-Driven Parametric Anomaly Synthesis (PDPAS) module employing multiple parametric functions to generate synthetic anomalies, providing explicit anomaly supervision. Then, we propose the Hierarchical Global-Local Anomaly Alignment (HiGLA) module to align global and local representations within the normal and anomalous groups. Finally, we propose the Semantic-Geometric Anomaly Classification (SGAC) module to jointly learn localization and classification, yielding spatially precise and type-discriminative anomaly representations. Extensive experiments establish new state-of-the-art performance on all four benchmarks. AT3D-AD achieves Object/Point AUROC scores of 98.1\%/98.9\% on Anomaly-ShapeNet and 95.0\%/95.2\% on Real3D-AD, while reaching 74.2\% Macro-F1 for anomaly-type recognition on Real3D-AD.

Tue 22 SeptComputer Vision and Pattern Recognition
The gist
Finding defects in 3D scans of objects is hard because it's often unclear exactly where the problem is and what kind it is. The authors created a new method called AT3D-AD that makes fake defects to learn from, then looks at the overall shape and small details together. It can spot, find, and name different kinds of problems better than earlier tools. Their tests show it works very well on several 3D datasets used for checking objects.
Open → 2609.25930v1

Multi hypersphere model improves anomaly detection with clear decisions

Interpretable Multi-Hypersphere Deep Anomaly Detection for Open-set Supervised Anomaly Detection

Abstract: Multi-class open-set anomaly detection requires a model to characterize the normal acceptance domain formed by multiple heterogeneous subdistributions using only class-labeled samples from known normal classes, and to identify previously unseen anomalies at test time. Existing single-hypersphere methods cannot explicitly represent class-specific locations and acceptance ranges, while current multi-hypersphere or multi-class approaches do not fully integrate inter-class boundary constraints, learnable acceptance ranges, and interpretable decisions. To address these limitations, we propose Interpretable Multi-Hypersphere Deep Anomaly Detection (IMHD-AD). IMHD-AD constructs an independent hypersphere for each known normal class in a shared feature space. With target-inside and non-target-outside constraints, IMHD-AD embeds the class-specific hypersphere centers and radii directly into the final network layer and jointly optimizes them with the shared representation. The minimum signed boundary score across hyperspheres simultaneously determines open-set acceptance or rejection and provides a faithful geometric explanation of each decision. On MNIST, Fashion-MNIST, and CIFAR-10, IMHD-AD achieves the highest AUC in 28 of 30 open-set comparisons. A two-dimensional synthetic study further shows that model architecture must balance the compactness of known normal classes against the separability of unknown anomalies.

Sat 19 SeptMachine LearningArtificial Intelligence
The gist
Detecting unusual data that doesn’t belong to any known category is hard when normal data comes in many types. The authors present a model that uses separate, interpretable shapes called hyperspheres for each normal type to better spot anomalies. This model learns where each normal category is and how big its range is, making it easier to decide if a new item fits or is an anomaly. Tests on popular image datasets show this method outperforms others in finding unexpected cases.
Open → 2609.23008v1

New method finds reliable minimum number of data distribution changes

Distribution-free inference on the number of changepoints

Abstract: Suppose we are given an ordered sequence of independent data whose distribution changes $K$ times at unknown locations, for some unknown $K \geq 0$. In this paper, we study the problem of performing distribution-free inference on $K$. First, we show an impossibility result: any distribution-free upper confidence bound on $K$ must be trivial and uninformative. Then, using conformal $p$-values, and under only the assumption that the data segments induced by the changepoints are exchangeable (within themselves) and mutually independent, we construct a finite-sample valid lower confidence bound on $K$, which we call the Conformal LOwer bound on Changepoint Count (CLOCC). We show that CLOCC is the only feasible way to provide a lower bound on $K$ under the stated assumptions, a property we refer to as its universality. We provide practical guidelines for choosing score functions that yield efficient and tight lower bounds. We evaluate CLOCC in several synthetic and real-data experiments, where it provides informative lower bounds on $K$, demonstrating its practical applicability.

Tue 8 SeptMachine Learning
The gist
Sometimes data changes its behavior at unknown points, and it's hard to tell how many changes happen. The authors show that you can’t confidently guess an upper limit for these changes without assuming something about the data. However, they created a new tool called CLOCC that can reliably provide a lower limit on how many changes occurred, based on minimal assumptions. CLOCC works by using a statistical technique called conformal p-values and gives useful results on both fake and real data.
Open → 2609.08234v1

Vision language models improve zero shot anomaly detection accuracy

Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection

Abstract: Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to object-centric bias, normal and anomalous text prototypes exhibit a high semantic overlap. While enforcing strict orthogonality between them improves discriminability, mapping highly contiguous visual inputs onto drastically orthogonal prototypes introduces a geometric dilemma, disrupting the pre-trained structural continuity. To address this problem, we propose Proximity-CLIP, a framework that visually calibrates the semantic margin to guide visual adaptation. First, we introduce a visually-calibrated semantic proximity learning mechanism that uses a bounded dynamic regularization to learn an appropriate semantic margin, ensuring discriminative separation while preserving structural alignment. Second, we design an Anomaly Query Module (AQM) driven by these text priors. Using the calibrated anomalous prototype as a semantic query, the AQM actively retrieves localized defect cues from contextual visual patches, mitigating the dilution of subtle anomalies during global pooling. Extensive experiments demonstrate that Proximity-CLIP outperforms current state-of-the-art methods across multiple ZSAD benchmarks with minimal architectural modifications.

Mon 7 SeptComputer Vision and Pattern Recognition
The gist
Detecting unusual or defective items in images without prior training is difficult because normal and abnormal features often look similar. The authors propose Proximity-CLIP, which adjusts how text descriptions relate to visual data to better separate normal from abnormal features. This method also focuses on small defective areas using a special querying mechanism and preserves smooth visual structure learned from before. Their approach performs better than existing methods on several tests with only minor changes to the model.
Open → 2609.07229v1