Image Caption Generator Using Neural Networks
Sujeet K. Shukla, Saurabh Dubey, Aniket Kumar Pandey, Vineet Mishra, Mayank Awasthi, Vinay Bhardwaj · International Journal of Scientific Research in Computer Science Engineering and Information Technology · 2021
In this paper, we focus on one of the visual recognition facets of computer vision, i.e. image captioning. This model’s goal is to come up with captions for an image. Using deep learning techniques, image captioning aims to generate captions for an image automatically. Initially, a Convolutional Neural Network is used to detect the objects in the image (InceptionV3). Recurrent Neural Networks (RNN) and Long Short Term Memory (LSTM) with attention mechanism are used to generate a syntactically and semantically correct caption for the image based on the detected objects. In our project, we're working with a traffic sign dataset that has been captioned using the process described above. This model is extremely useful for visually impaired people who need to cross roads safely.