A Text-to-Image Generation Method Based on Multiattention Depth Residual Generation Adversarial Network
Shuo Yu Yang, Xiaojun Bi, Jian Xiao, Jing Xia · 2021 7th International Conference on Computer and Communications (ICCC) · 2021
In this paper, we focus on generating realistic images from text descriptions. The current methods mainly use stacked networks, and there are two main problems. (1) The text information is relatively complicated, and it is difficult for the convolutional layer to extract deep-level text features. So the text features often can’t guide the generation of images effectively, which leads to a lower degree of matching between the generated image and the text; (2) The quality of the generated images is poor, and the authenticity, vividness and diversity of the images need to be improved. In this article, we propose a multi-attention depth residual generation adversarial network (MADR-GAN) based on the DM-GAN model to generate high-quality images. The network proposes a deep residual self-attention mechanism (DRSAM) in the low-resolution image generation stage, which is utilized to lift deep-level text information features and improve the quality of initial generated images. In the high-resolution image refinement stage, the CBAM attention mechanism is introduced to enhance the quality of high-resolution image generation. We evaluate the DM-GAN model on the Caltech-UCSD Birds 200 dataset. Experimental results demonstrate that our MADR-GAN model performs favorably against the state-of-the-art approaches.