CaptionCraft: VGG with LSTM for Image Insights

Syed Abudhagir U., Korpol Vignesh, Udugula Harish, Vuppala Avinash, M. Venkata Subbarao · 2024

In the dynamic field of AI, image captioning has emerged as a significant research area, leveraging advancements in deep learning. This study focuses on the synthesis of image descriptions through a meticulous process encompassing feature extraction and caption generation. Feature extraction is accomplished using the VGG-19 architecture with ImageNet weights, a powerful convolutional neural network renowned for its ability to capture intricate visual details. The model classifies objects within images, distinguishing between humans, animals, and plants. For generating captions, a Long Short-Term Memory (LSTM) network with 256 units is employed, facilitating sequential word-by-word prediction. The chosen architecture and weights contribute to effective feature representation, while the LSTM model excels in generating contextually coherent captions. The training process utilizes the Flickr 8k dataset, derived from ImageNet, providing a diverse set of images for comprehensive model training. Hence, the proposed model obtained a BLEU score of 0.669135. This holistic approach showcases the synergy between image processing and natural language processing, underscoring the model's proficiency in automatic image captioning.

Read the paper · More papers on PaperTik