Demosaicking Algorithm Using Swin Transformer and Long‐Range Attention Network
Jin Wang, Oh‐Jung Kwon, Gwanggil Jeon · Expert Systems · 2025
ABSTRACT Many mobile devices—including digital cameras, smartphones, and personal digital assistants (PDAs)—rely on single image sensors to capture scenes for real‐time processing. Convolutional neural networks (CNNs) have shown outstanding performance in various image processing tasks. In this paper, we propose a novel demosaicking method based on a transformer and long‐range attention network (TLAN). The approach begins by initializing the mosaicked image using a bicubic interpolation algorithm, which provides a coarse reconstruction. TLAN is then applied to refine the output and accurately reconstruct the three colour channels. Our TLAN architecture combines the Swin Transformer (ST) with a dedicated long‐range attention (LA) mechanism. The overall framework consists of both shallow and deep feature extraction modules. The deep extraction module is built from multiple residual swin transformer blocks (RSTBs), each composed of several Swin Transformer layers and augmented with a long‐range attention block (LAB) to capture extended spatial dependencies. Experimental results demonstrate that the proposed method achieves superior performance in both PSNR and visual quality compared to existing demosaicking techniques.