Text Guided Image Manipulation using LiT and StyleGAN2

Shantanu Todmal, Tanmoy Hazra · 2024

The convergence of generative adversarial networks and advanced natural language processing models in recent years has paved the way for crucial developments in text-based image manipulation. However, the conditional generative models rely heavily on the annotated text-image data for accomplishing this task. In addition, the quality of manipulation depends upon the accuracy of the textual descriptions present in the dataset. To address these challenges, we propose a novel framework for text-based image manipulation that eradicates the dependency on the text annotations for training. We leverage the characteristic of a well-aligned distribution between the differences in LiT's image features of two images and the differences in LiT's textual embeddings of two texts. The difference between the LiT encodings of the source and target image is mapped to the edit directions in the StyleGAN2 inversion space W+ using a mapper, during the training stage. The trained mapper then predicts the StyleGAN2 edit directions using the differences in the LiT's text encodings of the provided text prompts, during the inference phase. As a result, the mapper is trained without using text and generalizes well to a variety of text-based prompts. The proposed approach is able to achieve an FID score of 33.67 and IDS score of 71.89, better than the other state-of-art approaches like TediGAN, StyleCLIP, StyleMC, HairCLIP, CLIPInverter etc, which signifies superior image quality and precise manipulation.

Read the paper · More papers on PaperTik