Turning Images into Words: A Neural Approach to Image Captioning Using VGG16 and LSTM

K Sai Sarath Chandu, A Rakshith, V Nivethitha, Suthir Sriram · 2024

Recently, the task of automatic image captioning has been considered as a challenging and important problem for both computer vision and natural language processing domain. The primary aspect of image captioning is to create textual descriptions that will mirror the content of an image. This task has real significance in several real-world uses including automatic text writing, helping the visually impaired and improving image search technologies. In this paper we present a deep learning-based image captioning model using CNN and LSTM for generating accurate and fluent image description. In particular, we utilize the VGG16 model, one of the popular CNN architectures, for generation of deep features from input images. These feature vectors are subsequently fed into LSTM network as input captions are generated through understanding temporal structure in language.

Read the paper · More papers on PaperTik