Image Captioning and Identification of Dangerous Situations using Transfer Learning
Ruchika Malhotra, T.Pushpa raj, Vedika Gupta · 2022 6th International Conference on Computing Methodologies and Communication (ICCMC) · 2022
Expeditious advancements in Artificial Intelligence have drawn the recognition of numerous researchers towards Image Captioning and its applications. Clubbing the concepts of computer vision and natural language processing, image captioning has become an integral aspect of interpreting a scenario by generating captions. In this paper, a model is proposed which generates a caption as output for images as an input and tells if the image input is an image of the dangerous situation or not by the means of transfer learning. First, a dataset of thousand dangerous images are built and meaningful captions are linked with them. Dangerous situation here includes various potentially dangerous situation for humans, such as vehicle accidents, fire accidents, injuries, burns, murders, etc. The novelty of this work is enhanced by the dataset containing 1000 dangerous images and their self-created captions. The dangerous images contain photos of injuries, accidents, blood, murders and kidnapping. The proposed Image Captioning model uses ResNet50 for encoding of images, RNN and LSTM for sentence generation. Apart from that, transfer learning concepts have been included. The deep learning model could achieve accuracy of 70 on the given metrics, recall of 84, F1 score of 77.8 and Meteor score of 27.6 on a combined datasets of the thousand dangerous images datasets and the Flickr8K datasets. Specificity value of 65.1 and area under ROC curve value of 70 have been obtained.