Multimodal Learning for Image Caption Generation: A Deep Learning Approach
Prajakta Dhamanskar, Winster Pereira, Sharli Khot, Meetali Bhole · 2024
In today's highly visual digital landscape, images have emerged as a common mode of communication across social media platforms. However, attaching meaningful and contextually relevant captions to these images remains a significant challenge. This paper presents an automated image caption generator designed to overcome these hurdles by leveraging the power of deep neural networks. The proposed approach employs a hybrid architecture that combines convolutional neural networks (CNNs) and recurrent neural networks with long short-term memory (LSTM) units. CNNs, specifically the VGG16 and ResNet-50 models, are utilized for their proficiency in extracting salient visual features from images. The Transformer architecture, widely used in natural language processing, is compared against these CNNs for its attention-based mechanisms and language modelling capabilities. LSTMs facilitate the generation of coherent and linguistically accurate textual descriptions. The system demonstrates an excellent performance across standard evaluation metrics like BLEU, METEOR and ROUGE on the Flickr8K dataset. Extensive experiments conducted on the Flickr8K dataset validate the efficacy of the system in capturing the essence of diverse images and producing relevant captions. The modular design and straightforward training methodology enable convenient deployment across various platforms and applications. This work introduces an adaptable framework poised to revolutionize visual storytelling by bridging the gap between imagery and text in our increasingly interconnected visual world. The automated image caption generator represents a significant stride towards enhancing communication and understanding through the seamless fusion of visual and linguistic modalities.