Visual State Space Model for Image Super-Resolution

Jinhang Zhang, Min Gao, Wenzhao Li, Dan Fang, Chaowang Li · IEEE Transactions on Instrumentation and Measurement · 2024

Transformers and convolutional neural networks (CNNs) have garnered significant attention recently for low-level vision tasks, particularly image super-resolution (SR). However, CNNs are constrained by their local feature extraction capabilities, while Transformers suffer from the quadratic complexity of attention computation. To address these challenges effectively, we propose the dense-residual-connected mamba (DRCM) for SR. The DRCM overcomes the limitations of CNN and provides new advanced modeling capabilities akin to those of Transformers by utilizing global receptive fields and dynamic weighting mechanisms. We have integrated the Mamba with dense-residual connections to create a new dense-residual state-space block (DRSSB), which enhances local detail recovery and further reduces channel redundancy. DRCM effectively captures long-range dependencies, enhancing local detail recovery and significantly reducing channel redundancy. Such structure improves both the model’s efficiency and effectiveness. Experimental results demonstrate that the DRCM significantly improves the quality of reconstructed images while maintaining linear computational complexity, thereby achieving state-of-the-art performance on multiple benchmark datasets.

Read the paper · More papers on PaperTik