Multi‐Scale Feature Guided Transformer for Image Inpainting

Zeji Huang, Huanda Lu, Yu Xin, Hui Xiao · IET Image Processing · 2025

ABSTRACT In recent years, image restoration has witnessed remarkable advancements. However, reconstructing visually plausible textures while preserving global structural coherence remains a persistent challenge. Existing convolutional neural network (CNN)‐based approaches are inherently limited by their local receptive fields, often struggling to capture global structure. Previously proposed methods mostly focus on structural priors to address the limitation of CNN's receptive field, but we believe that texture priors are also critical factors that influence the quality of image inpainting. To tackle semantic inconsistency and texture blurriness in current methods, we introduce a novel multi‐stage restoration framework. Specifically, our architecture incorporates a dual‐stream U‐Net with attention mechanisms to extract multi‐scale features. The mixed attention‐gated feature fusion module exchanges and combines structure and texture features to generate multi‐scale fused feature maps, which are progressively merged into the decoder to guide the Transformer to generate more realistic images. Additionally, we propose a feature selection feedforward network to replace traditional MLPs in Transformer blocks for adaptive feature refinement. Extensive experiments on CelebA‐HQ and Paris StreetView datasets demonstrate superior performance both qualitatively and quantitatively compared to state‐of‐the‐art methods.

Read the paper · More papers on PaperTik