Masked Visual Transformer for Efficient Training with Small Dataset

Chen-zhi Guan · International Journal of Pattern Recognition and Artificial Intelligence · 2023

Vision Transformers (ViTs) are becoming an architectural paradigm replacing the convolutional neural networks (CNNs) in computer vision. ViTs offer competitive performances with respect to CNNs, but they, especially the vanilla ViTs, are hungrier for data than the typical CNNs as the short of the common inductive bias of convolution. Recently, few works focused on training vanilla ViTs efficiently with small datasets. In this paper, we perform research on training vanilla ViTs with small dataset containing thousands of images, and propose a method that is applied to self-supervised pretraining stage. The proposed method combines parametric instance discrimination with CutMix and Multi-crop. Furthermore, we introduce image masking to reduce the overfitting of pretraining on small dataset. State-of-the-art results are achieved by our method for training from scratch based on vanilla ViT backbones on seven small-scale datasets. The transferring performance of our method is also tested on small datasets, and results show that it is improved significantly.

Read the paper · More papers on PaperTik