Efficient Lightweight Super-Resolution with Swin Transformer and Attention Mechanisms
Ying Yu, Xuemei Sun, Tianxiang Liu, Zhen Tao Yu · 2024
In the field of image super-resolution, both Convolutional Neural Network (CNN) and Transformer based approaches have achieved impressive achievements. However, these methods are generally associated by high computing costs and a huge number of parameters. In order to solve these problems and achieve a better balance between the quality of image super-resolution reconstruction and network parameters, this paper proposes a lightweight image super-resolution network HATN based on Swin Transformer and attention mechanism, which combines the advantages of Swin Transformer and attention mechanism. HATN employs a$3 \times 3$convolution to extract shallow features of an image and a deep feature extraction module consisting of numerous residual attention Transformer blocks for deep image feature extraction. The shifted window self-attention mechanism of the Swin Transformer is utilized to capture the global contextual information of the image, but also by introducing an enhanced spatial attention (ESA) block and a contrast-aware channel attention (CCA) block, which further improves the network's ability to capture image details and structural features. The model is assessed on benchmark datasets such as set5 and set14. The experimental results show that the model proposed in this paper outperforms state-of-the-art lightweight super-resolution networks such as RFDN, ESRT, HNCT and BSNR on the lightweight super-resolution task, and achieves good performance on the Set5 dataset with an improvement of 0.02 dB compared to BSRN and 0.06 dB compared to HNCT, respectively.