Text to Face generation using Wasserstein stackGAN

Amit Kushwaha, Pasam Chanakya, Krishna Pratap Singh · 2022 IEEE 9th Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON) · 2022

Text-to-image generation is a difficult problem; with face datasets, the difficulty increases due to the finer details the models must learn. A model is proposed with the backbone as stack architecture with two stages. Stage-I generates low-resolution face images, which are then taken as input by Stage-II to generate high-resolution face images. Wasserstein loss has been used to stabilize training for both stages, and multiple text embeddings present for each real face dataset are provided while training in different epochs. In terms of FID score, the proposed model performs better than models like AttnGAN and DF-GAN and is comparable to TediGAN, which is the state-of-the-art architecture for text-to-face generation.

Read the paper · More papers on PaperTik