A Review on Automatic Image Caption Generation for Various Deep Learning Approaches
Arun Pratap Singh, Manish Manoria, Sunil Joshi · 2023
In machine learning, image captioning is a process of extracting information as per the action performed in an image. An image may contain various objects that may be associated with certain actions; image captioning is responsible to translate the action into text by the help of deep learning algorithms. Image captioning is a challenging task in the field of image processing because real world data is bit complex in action and may contain unidentified objects. It has been implemented in various deep learning algorithms such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), DNN (Deep Neural Network), LSTM (Long Short Term Memory) and many more. These networks use distinct datasets such as Flickr8K, Flickr30K, MSCOCO and many more to obtain the result accordingly. Flickr8K contains eight thousand images for training; similarly Flickr30K contains thirty thousand images and MSCOCO contains 82783 images that is why MSCOCO is considered as large dataset and it is preferred by various researchers. The result of the machine translation is based on syntactic and semantic traits. The quality of machine translation can be measured by BLEU (BiLingual Evaluation Understudy) in the reference of ground truth captions. The intension of this paper is to review various approaches related to the image captioning and it also intended to find out the best deep neural network model among various implemented models.