Image Captioning with CNN and LSTM using Python
Srikanth Bethu, N. Subhash Chandra · 2023
Our vision is our most vital sense. Software developers have utilized the capability of vision as they build more interactive, intelligent, and accessible software through images. However, there are scenarios where an image might not be sufficient alone. There might be extra context needed, or alternative text is displayed to circumvent bandwidth restrictions and provide a more accessible experience. In an era where there are simply a huge number of images to be described, the manual description fails at such a scale. Through the help of deep learning, image processing and natural language processing can combine to give way to empower the computer to describe images on its own. This can be presented as a consumable service through the use of a web frontend where users can simply provide the images they wish to be described. This allows anyone to leverage the power of this deep learning approach with relative ease and the use of adapti ve image descriptor functionality via an easy-to-use API while the computationally heavy tasks will be abstracted away. In this paper LSTM (Long Short-Term Memory) and RNN (Recurrent Neural Networks) are discussed.