Text-to-Clipart using AttnGAN

Yusuke Kondo, Takafumi Sakura, Toshihiko Yamasaki · 2020

The Attentional Generative Adversarial Network (AttnGAN) [1] is a state-of-the-art text-to-image generation model. One of the key factors of AttnGAN's success is the ability to evaluate the similarity between an input sentence and the generated image in the same feature space (Deep Attentional Multimodal Similarity Model, DAMSM). However, the network architecture of AttnGAN is complicated and vast, which necessitates considerable computational costs in the training process. When AttnGAN is applied to a text-to-image generation task in different image domains such as clipart, the output images are simpler than the high-resolution natural images that AttnGAN originally assumes. Therefore, we propose a lightweight AttnGAN aiming at reducing the training computational cost without compromising the quality of the generated images. In particular, we focus on the image encoder; replacing it from Inception-v3 to VGG-16 reduces the DAMSM training time by approximately half of the original implementation.

Read the paper · More papers on PaperTik