Image Captioning Using Deep Learning Techniques for Partially Impaired People
Suganthe Ravi Chandaran, Shanthi Natesan, Geetha Muthusamy, Pravin Kumar A L Sivakumar, P. Mohanraj, Richard Joyal Gnanaprakasam · 2023
Image captioning is used to describe an image based on the features and actions that are present in that image. The existing image captioning primarily uses an encoding and decoding structure, with the encoder which extracts image features using CNN model as well as the decoder uses LSTM model. The current encoding and decoding scheme makes extensive use of the attention mechanism. Gradient explosion is a problem with the current image caption models, which are based on recurrent and convolutional neural networks, and therefore are not particularly good at extracting valuable information from images. To overcome this problem, the YOLOv5 and Bidirectional LSTM model is proposed. YOLOv5 is used to identify the objects in a given image and a bidirectional LSTM (Bi-LSTM) layer is used to extract the features of the given image. This proposed algorithm gives the optimized result with good accuracy. The Flickr8k benchmark dataset is used to test this approach. From the results it is known that the trained model performs better than alternative encoder-decoder methods that depend solely on global image features. The metric which has been used to evaluate the model is BLEU which is mainly used for machine translated text evaluation. This model gave a 0.7 BLEU score.