Enhanced Image Captioning for Social Media using Inception V3 and Transformer Networks

T. Maheshwaran, K Ragul, Monish Coumar S, A Narendheran · INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2024

In this digital age, social media is flooded with tons of visual content, automatic image captioning is a must for better content accessibility and engagement. This project is an advanced approach to image captioning using Inception V3 for feature extraction and Transformer for natural language generation. Inception V3 is known for its object detection capabilities, it serves as the backbone to capture the finedetails of the image while Transformer with its attention mechanism generates contextually rich and coherent captions. Our system aims to provide more accurate and context sensitive descriptions for social media images, to improve user experience and content discoverability. We used Flickr8k dataset for training and testing, we show the effectiveness of this hybrid model in handling complex scenes, multiple objects and varying context. The proposed solution is a scalable and efficient way to generatecaptions that will enhance social media interaction and accessibility. Keywords: Image Captioning, Inception V3, Transfomer Network.

Read the paper · More papers on PaperTik