Stacked Generative Adversarial Networks (StackGAN) Text-to-Image Generator
Rahul Arya, Dikshant Joshi, Khushi Sharma, Vihan Singh Bhakuni, Satvik Vats, Vikrant Sharma · 2024
Generating high quality images from textual description is an ongoing challenge which can have a wide range of practical uses. Some of such fields are computer vision, creation industry, e-commerce, education and many more. Currently existing text to image synthesizing techniques, while offering valuable insights still struggles to capture fine details and vivid object components that are crucial for real world application. We use an approach that uses Stacked Generative Adversarial Networks (StackGAN) to solve the problem of text-to-image generation. This includes using a two-stage process. In the first stage (Stage – 1 GAN), we aim to develop a system that is capable of generating the fundamental shape and color attributes of the object according to the text. In this stage, a low-resolution image is generation that is used as a foundation for the subsequent image refinement process. In the second stage (Stage – 2 GAN), we focus on enhancing the quality of image by incorporating the text descriptions and generating high-resolution image. This project aims to bridge the gap and advance the text-to-image generation to create images to create images that are textually accurate. By doing so, the project seeks to unlock the potential of this technology across various fields and bridge the gap between the power of language and art, design and content creation.