Deep Learning based Image Captioning Models: A Critical Analysis
Himanshu Sharma, Aniket Singh, Gouri Shrivastava · 2021
The main goal of image captioning is to produce moat appropriate caption for any image. It is required to find out objects in a picture, their connection and also some hidden features which might be not present in a picture. When the recognition is done, our succeeding step is to produce the most suitable, accurate and short description for a picture which must be grammatically correct. Computer vision concept is used for recognition of objects and description is generated by methods of natural image processing. Sometimes it is not easy for a machine to emulate exactly like human brain or thinking like that. However, in this field researches have shown their great achievements. By using deep learning algorithms like CNN and RNN, we are capable enough to deal with such type of problems. In intelligent control system and IOT based devices this is mostly used. In this research paper, we are introducing distinct models of picture description like retrieval based, template based or deep learning based and distinct evaluation methods also.