Waveformer: Innovative Token Mixing for Plug-and-Play Image Restoration
Dingjie Peng, Wataru Kameyama · 2024
Recent advancements in Transformer-based methods have demonstrated remarkable performance in various low-level vision tasks, including image denoising, image deblurring, and image demosaicing. However, many of these methods face limitations in plug-and-play image restoration due to two key drawbacks: inefficiencies in the Self-Attention mechanism of Transformers and the inability to dynamically adjust normalized features using spatial information from the input image. To address these issues, we propose a novel attention mechanism called Wave-Attention and use it as a token mixing strategy to build Waveformer, an asymmetric U-Net model. In the encoder, Wave-Attention aggregates tokens in wave representation, efficiently capturing dependencies based on learned amplitude and phase. Additionally, our decoder utilizes a revised spatial adaptive normalization (Spade) module to enhance normalized features using the spatial context of the input noisy image. The proposed encoder-decoder architecture underpins an effective deep learning-based denoiser. In our experiments, we first train Waveformer using noise level maps and noisy images as inputs for image denoising, enabling the model to handle a wide range of noise levels. We then integrate the well-trained model into the half quadratic splitting algorithm as a denoiser to tackle image deblurring and image demosaicing tasks. Our experimental results demonstrate that the proposed method surpasses the latest learning-based methods in various image restoration tasks.