Automatic Synthesis of Realistic Images From Text using DC-Generative Adversarial Network (DCGAN)

K. Deepthi, K. Aditya Shastry · 2023

Automatic Conversion of text to realistic images is an interesting, amazing and useful technique. There are several methodologies proposed by many researchers in achieving this but none of them have addressed this effectively. They lack efficiency in terms of images generated in terms of quality or precision. This work intends to generate realistic and semantically correct images from the given input text. Generative Adversarial Networks (GAN) are used to convert human written description text into its corresponding image pixels. The model is designed to initially learn the features hidden in text to capture the unique visual details. These factors are further used to synthesize semantically correct images which humans might mistake for real images. The data set, the Oxford 102 flower dataset is used which has at least five text descriptions describing the image. The specific Deep Convolutional - Generative Adversarial Network consisting of a Generator and a Discriminator is used to implement the proposed model. DC-GAN is a direct extension of the GAN which explicitly uses convolutional and convolutional-transpose layers in the discriminator and generator. The aim is to train the DC-GAN model to produce compelling and accurate images for the corresponding input description. The clarity of the image obtained is directly proportional to the number of epochs the model was trained for. The greater the number of epochs the model is trained, the better the efficiency/clarity of the image. The model is evaluated for the various performance metrics. The model is evaluated based on the image clarity obtained and the context of image generation. The performance of the model is measured based on the clarity of the images generated and the context for which the image was generated.

Read the paper · More papers on PaperTik