Learning Visual Augmentations from Linguistic Knowledge Prior for Anomaly Detection

Xiaohan Jiang, Xinpeng Liu, Zhiqiang Ding, Xi‐Ming Sun, Zhenjiang Hu · 2025

With the advancement of large vision-language models (VLMs), significant progress has been made in anomaly detection. Among these models, CLIP, renowned for its exceptional multi-modal representation capability, has demonstrated outstanding performance on various downstream learning tasks. However, CLIP primarily aligns paired text prompts with global image-level representations, limiting its ability to capture finegrained alignments between image regions and text spans, thereby constraining its effectiveness in precise visual anomaly localization. To address this limitation, we propose a two-stage framework that enhances the capablity of CLIP for anomaly detection. In the first stage, we propose trainable anomaly feature augmentations to learn linguistic knowledge from prompts. Meanwhile, image-level and patch-level alignments are introduced to enhance the detector's ability to identify anomalous objects across diverse contexts. In the second stage, the detector is finetuned base on the learned augmentations. Specially, normal object features are transformed into pesudo anomaly object features via augmentations, enabling anomaly object detection. Finally, we validate the effectiveness of our approach on the MVTecAD dataset, from which the superior performance in anomaly detection and localization can be demonstrated.

Read the paper · More papers on PaperTik