Image captioning System Using LSTM and VGG16

Manish Kumar Singh, Santosh Kumar Upadhyay, Nidhi Sharma, Rupak Kumar, Manmohan Singh Yadav · 2024

The process of creating literary descriptions or captions for images based on their content is referred to as automatic picture captioning. This machine-learning task combines natural language processing (NLP) and computer vision to generate the captions. Currently, auto photo captioning is a highly advanced and rapidly evolving research field, with new methods being proposed regularly to improve results. However, there remain challenges to achieving human-level performance. This analysis aims to build on the findings from previous research to push the boundaries further. In this study, we extracted features from images using the VGG16 model of CNN, which serves as an encoder, along with the Flickr dataset. A Long Short-Term Memory acts as a decoder, using the built-in lexicon and image features to generate captions for the photographs. The extracted image features are converted into a feature vector, which is then input to the LSTM to create representations that correspond to the visual content of the image. The captions describe various aspects of the image, such as the object's name, location, colour, dimensions, features, and surroundings. BLEU (1 to 4) is the most used evaluation metric in all studies. The LSTM algorithm combined with CNN has also been found to outperform RNN combined with CNN. We identified that the encoder-decoder architecture and the attention mechanism are two of the most promising methods for implementing this model. Utilizing both approaches together can significantly advance the project.

Read the paper · More papers on PaperTik