Unsupervised Image translation model based on Transformer and CycleGAN
Yuhang Cao, Hong Ye Yang, Fei Shang · 2024
This paper addresses the challenge of unsupervised image translation, which aims to transform images from the source domain into the context of the target domain. Current approaches often fall short in producing images with realistic textures, preserving input fidelity, and achieving robust model generalization. In order to get over these restrictions, we introduce a composite model that synergistically combines the Vision Transformer (ViT) with CycleGAN. Our innovative approach integrates a Depthwise Separable Convolution-Multilayer Perceptron (DWconv-MLP) module to effectively capture rich contextual information. The model further incorporates moving average strategies and spectral normalization within the discriminator network to enhance performance. To mitigate overfitting and alleviate the issue of vanishing gradients, we introduce an improved zero-centered gradient penalty method. The efficacy of our model is demonstrated through rigorous experimental evaluation on well-established image translation datasets, showcasing its superior capability in style feature extraction and production of visually compelling images that maintain a strong resemblance to their original counterparts.