AplusN: Progressively Integrating Attention and Normalization in Wavelet Domain for Pose Transfer
Wei Yu, Rui Wang, Weizhi Yang, Wenjian Hu, Wei Xiang · IEEE Transactions on Multimedia · 2025
Pose-guided person image generation aims to synthesize images of human in various poses, often encountering issues such as occlusions and texture transfers. Previous methods have utilized attention mechanisms, flow field, normalization techniques, and diffusion model. Among them, flow field and attention are the two most commonly used methods. Flow fields are good at preserving detailed textures, while attention is better at generating reasonable semantic structures. Previous networks often used only one of the two and failed to make full use of their advantages. At the same time, the flow field and attention also showed complementary functions in the frequency domain. The flow field was good at preserving the high frequency information of the image with the detailed texture, while the semantic structure of attention was good at generating the image with the low frequency information, and few networks used this to improve the generation effect. Based on these facts, this paper introduces the AplusN network, which innovatively addresses the image generation problem by processing from low to high frequencies. For low-frequency information, a conditional large-kernel convolutional attention mechanism (CLA) is employed to capture the global information of the human body. High-frequency information is refined using a spatial-channel normalization module (SCN) to enhance the body's detailed textures. Additionally, we propose a wavelet loss function to align the frequency domain information of the generated images with the target images. Both qualitative and quantitative experiments demonstrate the superiority of our method over state-of-the-art (SOTA) methods, yielding better-defined overall body contours, local details, and higher-quality image generation.