OV-AS: Zero-Shot/Few-Shot Open-Vocabulary Anomaly Segmentation Based on CLIP

Jingtong Mo, Yuzhuo Fu, Ting Liu · 2024

Existing anomaly detection methods mainly focus on unsupervised learning, resulting in low generalization and single type segmentation. Hence, we introduce open-vocabulary detection pattern into anomaly segmentation field to achieve multi-semantic segmentation and propose a zero-shot/few-shot open-vocabulary anomaly segmentation OV-AS based on CLIP. (1) In zero-shot, introduce pretraining on the open-vocabulary segmentation models for specific anomaly domain knowledge learning based on an extracted Anomaly Domain-General dataset. Then, introduce fine-tuning on the image encoder of CLIP for further narrowing the gap with target domain. (2) In few-shot, introduce few-anomaly-shot learning on target datasets. Results show that in zero-shot, our model achieves 35.7/13.3 F1-max on MVTec-AD/VisA benchmarks, essentially surpassing state-of-the-art. In few-shot, our model achieves 43.0 mIoU on VISION benchmark under 10-shot, comparable to supervised learning.

Read the paper · More papers on PaperTik