An Analysis of Image Captioning Models using Deep Learning

Navya Goel, Aditi Arora, Priyanshu Kashyap, Sagar Varshney · 2023

Developing a self-descriptive software for describing and explaining the activities seen in an image has become a necessary and interesting field of research in Deep Learning and Artificial Intelligence. The process of formulating precise description for an image that is syntactically as well as semantically correct is known as Captioning. The task commences with classifying the attributes of an image and extracting feature vectors using various CNN models. In the proposed work, we have demonstrated the working of three different CNN models i.e. Xception, VGG-16 and ResNet50 along with observing the impressive accuracy achieved by each model respectively. The dataset used in our proposed work is Flickr_8k dataset having 8091 images. The features extracted from these three models individually are fed to LSTM (Long Short Term Memory) model, an example of RNN model and further using this for sentence formation. We have generated captions using each model and verified the accuracy achieved using visual metric and BLEU (BiLingual Evaluation Understanding) metric. The three systems grant the good quality result and captions, on comparing the BLEU scores which are, 0.79 for Xception Model, 0.75 for VGG-16 and, 0.84 for ResNet50. ResNet50 is proven to be best for classification and feature extraction with 84% accurate captions on 50 epochs and also facilitates the benefit of overcoming the vanishing gradient problem. We believe this technology can help visually impaired people to recognize and interpret the correlation of activities happening in the image.

Read the paper · More papers on PaperTik