Towards the Efficiency of the Fusion Step in Language-Based Fashion Image Editing
Mahta Darvish, Jamshid Shanbehzadeh, Azadeh Mansouri · 2022
Text-to-image synthesis is a new research field of image generation. The generated images must be based on textual descriptions and of acceptable quality. Among the generative models, Generative Adversarial Networks (GANs) can generate higher quality images compared to other models. Most GAN-based text-to-image generation methods use simple datasets, including flower and bird images. For more complex datasets (such as the Fashion Synthesis dataset used in this paper), the issue of generating an image from a text becomes more challenging, as the images in this dataset are much richer in content than just the flower images or birds. Most of the methods proposed so far are only able to recognize the color of an object based on the text description and have difficulty in accurately identifying the location in the image where these changes are to be made. One way to solve this problem is to generate an image from the text based on editing this image. In this paper, the methods of combining text and image features in the generation of images are discussed and the effect of different fusing models on the process in terms of quality and accuracy of the description is investigated.