SRViT: A Vision Transformer via Super-Resolution
Wenxuan Song, Zhenyu Wang, Yunjia Gao, He Xu · 2023
Vision Transformer is a widely used model in computer vision that can extract and learn useful feature information from images. However, its performance may be limited by low-quality and low-resolution datasets with complex degradation. Therefore, how to improve the accuracy of image classification is a challenging problem. One possible method is to increase the resolution of the input images or the input size of the network model. Both of these can improve the quality of the input, but only the latter has been verified experimentally. Ideally, a high-performance super-resolution network can enhance the resolution of the images and restore the details of the real-world images. To this end, we can use a super-resolution network to increase the resolution of the input images for Vision Transformer. In this paper, we propose a network structure that combines a super-resolution network and a Vision Transformer. First, the input image datasets are fed into a resolution classifier, which is a lightweight CNN and able to choose an appropriate resolution for each images. In this stage, a Gumbel softmax is inferred to keep differentiable. Second, the leveled images are sent to the super-resolution network, and blind super-resolution is performed at different magnification factors according to the level. Finally, the images are fed into Vision Transformer for image classification.