Deep learning based automatic image caption generation for visually impaired people

Pranesh Gupta, Nitish Katal · 2023

In conversational artificial intelligence, automated generation of the captions – that is, trying to explain the content and context of an image by means of well-molded sentences – is considered a strenuous affair. Unlike object detection, the image caption generation is far more challenging as it requires more relationships in the data, like trying to recognize the actions in the images, make a meaningful inference, and then generate the conversational sentences or correct descriptions for an image. This functionary can be used to assist a visually impaired person, where by using the camera sensor, the person can get an accurate description of their surroundings. Using this, a visually impaired person will be able to contextually under their surroundings in a much more natural and conversational way. The present work proposes the use of deep learning to solve this problem. The work proposes a hybrid system of CNN-RNN, where a multilayer CNN is employed to get the attributes from the images, followed by inputting them into Long Short-Term Memory (LSTM) to create meaningful captions for an image using the vocabulary created from the training network in the English language. In the work, Flicker8k dataset has been considered. The Bilingual Evaluation Understudy Score (BLEU) scores is used to evaluate the performance of the translated text by comparison with one or more reference sentences. From the obtained results, it has been observed that the Xception network offers good performance when compared to VGG16 model.

Read the paper · More papers on PaperTik