QwinSR: A Simplified MLP-Based Super-Resolution Model Integrating Swin-Mixer and SwinIR Architectures
Vimlesh Kumar · 2025
Despite advances in image processing, resolution loss remains a common challenge, creating a demand for superresolution (SR) techniques that reconstruct high-fidelity images from lower-resolution sources. This paper introduces QwinSR, a novel hybrid model for single-image super-resolution, which leverages the shifted window approach from Swin Transformer and the all-MLP design philosophy. QwinSR combines the shallow feature extraction capabilities of convolutional layers with the deep feature extraction of Swin-Mixer blocks, a derivative of the Swin Transformer that substitutes self-attention with multilayer perceptrons (MLPs). While focusing on simplicity and efficiency, QwinSR aims to achieve competitive performance with state-of-the-art SR models by utilizing residual connections and convolutional layers in its architecture. Our proposed model is evaluated through a series of ablation studies, where we analyze the impact of various hyperparameters, such as channel numbers, patch sizes, and the presence of residual connections, on the peak signal-to-noise ratio (PSNR) of upscaled images. Due to resource constraints, a full experimental benchmark was not conducted; however, the potential of QwinSR is discussed in relation to existing benchmarks in SR literature. We conclude by outlining future work, including formal testing on larger GPU clusters and further architectural enhancements, to refine the performance and efficiency of QwinSR in real-world applications.