IAD-CLIP: Vision-Language Models for Zero-Shot Industrial Anomaly Detection

Zhuo Li, Yifei Ge, Qi Li, Lin Meng · 2024

This paper presents an efficient zero-shot industrial anomaly detection (IAD) framework based on visual-language models. Industrial anomaly detection usually adopts an unsupervised learning approach, which achieves excellent detection performance though. However, it is still difficult to recognize some more complicated anomalies, such as rotational defects. At this point, more detailed features are needed to describe the image. With the excellent performance of contrastive language-image pretraining (CLIP), this paper proposes a zero-shot industrial anomaly detection framework IAD-CLIP based on visual language models. The framework contains a pre-trained CLIP model, a training-free adaptation module and a test-time adaptation mechanism. The training-free adaptation module uses a value-value attention mechanism and a state prompt space. The pre-trained CLIP model is used for feature extraction and the training-free adaptation module processes the extracted features through visual coders and text encoders for anomaly detection and localization. A test-time adaptation mechanism is used to improve the anomaly localization performance during the testing phase. The experimental results on the industrial anomaly detection dataset MVTec AD show that IAD-CLIP achieves 92.1% AUROC, 94.6% AUPR, and 91.9% F1Max, respectively. This result validates the significant effect of the IAD-CLIP framework proposed in this paper in the industrial anomaly detection task.

Read the paper · More papers on PaperTik