Image Captioning based on Deep Convolutional Neural Networks and LSTM
Swati Srivastava, Himanshu Sharma, Pragati Dixit · 2022
Image captioning is a challenging task that needs the knowledge from both computer vision algorithms and language processing techniques. The model must be able to understand an image and then apply language generation techniques to describe an image in a natural language such as English. In this paper, we have presented an image captioning model which uses VGG16 for visual feature extraction and LSTM model to generate sentences corresponding to extracted visual features. We have performed experiments on Flickr8k and Flickr30k datasets. Bilingual Evaluation Understudy (BLEU) metric is used to measure the accuracy of the proposed model. The proposed model can be further extended to wide range of applications related to IOT based applications and smart control systems.