Hybrid Feature and Sequence Extractor based Deep Learning Model for Image Caption Generation

Rohit Kushwaha, Anupam Biswas · 2021

Generating the captions from the input images is a crucial problem, it includes both computer vision and natural language processing. In this paper, we introduced a hybrid feature and sequence extractor-based deep learning model that helps in generating captions for the input images. The proposed generative model is designed using a deep convolution neural network (VGG-19) for extracting the salient feature vectors from the images. While producing the informative captions from these feature vectors, the LSTM (Long Short Term Memory) network is utilized. The model is trained to expand the probability of the target description sentence delivered to the training image. In the analysis, the model is trained with very renowned data sets like FLICKR-8K, and the correctness of the model is calculated utilizing the BLEU score granting the range 0 to 100. We compare our model based on the BLEU score with four other models.

Read the paper · More papers on PaperTik