Image Captioning Using CNN and RNN
S Rohitharun, L Uday Kumar Reddy, S Sujana · 2022 2nd Asian Conference on Innovation in Technology (ASIANCON) · 2022
Image captioning is a new task that has recently gained a lot of attention. It has many potential real life applications like virtual assistant, intruder alert in cctv functions and many more. A person can call out and explain an infinite number of elements of a visual situation with a quick glimpse. But automatically describing what’s in a photograph or image has always been a difficult task in Artificial Intelligence. Here the purpose is to summarise whatever is presented in the image in a single phrase, such as the items involved, their attributes, the activities being done, the interaction between the elements, and so on. The implementation of an Automatic Caption Generator employing CNN and RNN-LSTM models is described in this work. It integrates contemporary machine translation and computer vision research. Flickr8k was utilised as a dataset.