Improving Few-Shot Learning with Vision Transformer
Di Qi · Journal of Physics Conference Series · 2022
Abstract Vision Transformer (ViT) is emerging as an alternative to convolutional neural network (CNN) for visual recognition and has achieved impressive results, however, due to its data-hungry nature, ViT encoder in few-shot setting remains rarely explored. In this paper, we first propose a simple yet effective baseline that exploits the complementarity of self-supervised learning (SSL) and ViT, enabling the use of standard ViT model as the few-shot learner. Second, based on the baseline, we introduce a novel regularized fine-tuning framework, where the Parametric Instance discrimination (PID) and Base-Novel (BN) regularizations are proposed to reduce the intra-class variance and calibrate the biased distribution of novel classes to further enhance few-shot recognition. We conduct extensive experiments and show that our method can achieve new state-of-the-art performances on two widely used benchmarks.