Generate Detailed Captions of an Image using Deep Learning
Khan Shayaan Shakeel, Masalawala Murtaza Shabbir, Qazi Faizan Ahmed, Nafisa Mapari · International Journal for Research in Applied Science and Engineering Technology · 2022
Abstract: This paper shows the implementation of image caption generation using deep learning algorithm. The project is one of the primary example of computer vision. The main aim of computer vision is scene undertanding. The algorithms used in this model are CNN and LSTM. This model is an extension of the model based on CNN - RNN Model which suffers from the drawback of vanishing gradient. Xception model is used for image feature extraction and is a CNN model that is trained using ImageNet dataset. Extracted features from the Xception model is fed as the input to the LSTM model which in turn generates the caption for the image. The dataset used for training and testing is Flickr_8k dataset.