Transformer Encoder-Decoder Mask Reconstruction in Industrial Image Anomaly Localization

Xuhong Luo, Yongchun Liu, Guoming Chu · 2024

Acquiring a substantial amount of high-quality data for industrial image detection poses significant challenges in the field of computer vision. The imbalance between normal and anomalous samples, where normal samples far outnumber anomalies, complicates the training of traditional supervised detection methods, limiting their ability to exhibit robust detection performance on industrial image datasets. Unsupervised anomaly detection methods based on image reconstruction train exclusively on normal samples. These methods employ encoders and decoders to map and reconstruct images, assessing deviations from the distribution by evaluating unseen instances of both normal and anomalous cases. We propose a method for anomaly localization in industrial images using Transformer Encoder-Decoder Mask Reconstruction. The self-attention mechanism of the Transformer enables better attention to different positions within the entire input image in both the encoder and decoder, capturing long-range dependencies in the image and more accurately learning crucial features. Additionally, to train a stable and efficient image reconstruction network, we introduce a block-wise reconstruction approach involving masking the input images. To validate the algorithm's generality in industrial images, experiments are conducted on the MVTecAD and a custom crankshaft bearing datasets. Experimental results demonstrate the algorithm's excellent detection performance, surpassing previous state-of-the-art methods and confirming its effectiveness.

Read the paper · More papers on PaperTik