Adversarial Network for Photographic Image Synthesis from Fine-grained Captions
Padmashree Desai, C.N. Sujatha, Ramnath Shanbhag, Raghavendra Gotur, Rajesh Hebbar, Praneet Kurtkoti · 2021 International Conference on Intelligent Technologies (CONIT) · 2021
Automatic synthesis of realistic images from a given caption of text is a fascinating research idea and useful in many applications such as image in-painting, photo-editing, computer-aided design, etc. Current AI techniques are still exploring in this direction. However, researchers are currently developing text-to-image synthesis networks focused on the learning of discriminative text and image features using continuous and generalized robust neural network architectures. Many applications use deep convolution Generative Adversarial Networks (GANs) to construct extremely persuasive representations of explicit categories like faces, album covers, birds, creatures, and room interiors, among others. With the progress of generative models, neural networks can not only recognize images but are used to generate audio and realistic images as well. In the proposed work, the authors have used GAN-CLS architecture to create images from given text descriptions/captions. The experiment uses the CUB-200 dataset, which contains 11,788 bird images from 200 categories, as well as the Oxford-102 dataset, which contains 8,189 flower images from 102 categories. The proposed system's performance is assessed and compared to that of other systems.