Content-aware Fine-grained Sparse Transformer for Single Image Super-resolution
Qingtang Ding, Liangcheng Qin, Jungang Yang · 2023
Single image super-resolution (SR) is a classical low-level vision task, which has been investigated for many years. With the development of deep learning (DL), single image SR has achieved the promising performance in the past decade. Nowadays, DL-based single image SR methods roughly can be divided into CNN-based SR and Transformer-based SR. Current Transformer-based single image SR methods can model the long-range dependencies and capture the no-local self-similarity inside image over the CNN-based ones. Therefore, Transformer-based single image SR methods achieve the state-of-the-art performance. However, Transformer-based SR methods generally sample image tokens densely and calculate the multi-head self-attention (MSA) among all the tokens, which has the large computation costs. In this paper, we combine the sparsity nature of SR task and the powerful long-range modeling capability of Transformer to improve the efficiency and scalability of the Transformer-based SR methods. Specifically, we propose a content-aware fine-grained sparse Transformer (CFST) network for single image SR. Our CFST can efficiently calculate the MSA by clustering the related tokens in content and reduce the unnecessary computation between irrelated tokens. Extensive experiments show that our CFST can achieve the superior SR performance with less computational costs on SR benchmark datasets.