Text to image generation combined with generated caption using optimization techniques

B. Lakshmi Sowmya, Ramasamy Dhanalakshmi, B. Sai Teja, R. Nithin, S. Varshitha, S. Kalaivani · 2024

In the modern era, images serve as a primary means of communication and artistic expression. While humans effortlessly interpret visual scenes and articulate them in precise language, replicating this capability in machines poses a formidable challenge. Text-to-image synthesis is required in order to produce an image from a text-based description; this process is intricate. Similar to this, the goal of image to text captioning is to use natural language generation to convey the content of an appearance. Natural language processing and computer vision techniques are combined to create both text-to-image synthesis and image captioning. The application of techniques to images in order to enhance image quality or extract relevant information is the wide and diverse field of image processing. One effective strategy in these approaches is image composition. An image that bridges the gap between appearance and text by combining linguistic and visual information is referred to as a texted image. This helps viewers better comprehend the content. Therefore, this work presents a novel method that combines text-image synthesis with image captioning and image processing. The framework aims to improve the interpretability of images and the communicative capacity between images and texts. It consists of three primary components: an image combiner, an image captioner, and a text-to-image generator. Applications for the framework can be found in social networking, e-commerce, and education. The framework performance on several text-to-image datasets shows that it can produce a variety of high-quality images with appropriate and relevant captions. As a result, the new framework will enhance user experience and facilitate integrated, in-depth image analysis. In essence, this paper pioneers an integrated approach to text-image synthesis and captioning, leveraging image processing techniques to augment the comprehensibility and communicative potential of images. The framework success shows that texted images are one way of communication.

Read the paper · More papers on PaperTik