TT2INet: Text to Photo-realistic Image Synthesis with Transformer as Text Encoder

Jianwei Zhu, Zhixin Li, Huifang Ma · 2021

A text-to-image (T2I) generation method is mainly evaluated from two aspects, one is the quality and diversity of the generated images, and the other is the semantic consistency between the generated images and the input sentences. The feature extraction of the text is a very important part. In this paper, we propose a Transformer based Text-to-Image Network (TT2INet). we use the pre-trained Transformer model (ALBERT) to extract the sentence feature vectors and word feature vectors of the input sentences as the basis for the Generative Adversarial Networks (GANs) to generate images. In addition, we also added self-attention mechanism and spectral normalization method to the model. Adding a self-attention mechanism can make the model pay attention to more local features when generating images. Using the spectral normalization method can make the training of GANs more stable. The Inception Scores of our method on Oxford-102, CUB and COCO datasets are 3.90, 4.89 and 26.53, and R-precision scores are 92.55, 87.72 and 92.29, respectively.

Read the paper · More papers on PaperTik