MaskViT: Masked Vision Transformer for Multi-Class Unsupervised Anomaly Detection
Yunxin Liu, Danqing Liu, Yanhui Guo, Guanliang Wan, Min Zheng, Tengyue Yang · 2025
Unsupervised anomaly detection has become a critical task in industrial visual inspection. However, many reconstruction-based approaches suffer from the ”identical shortcut” problem, where models tend to simply replicate the input, thus limiting their anomaly detection performance. To address this issue, we introduce MaskViT, a novel framework for multi-class unsupervised anomaly detection. The method incorporates stage-specific masking strategies during the reconstruction process, which are tailored to the hierarchical characteristics of features, thereby enhancing both the semantic quality and sensitivity of the reconstructed representations. Furthermore, we design the Multi-Stage Channel Attention Fusion (MSCAF) module, which effectively aggregates multilevel features extracted by the encoder along the channel dimension, improving the discriminative power of the feature representations. Extensive experiments on the MVTec-AD and VisA datasets demonstrate that MaskViT achieves state-of-theart performance in both image-level and pixel-level anomaly detection tasks.