An Image Captioner for the Visually Challenged

Ariane Correa, Kaustubh Shetty, Reyna Binny, Ashwini Pansare · 2021 2nd Global Conference for Advancement in Technology (GCAT) · 2021

Spontaneously detecting and delineating the components of images is a core issue in the field of Artificial Intelligence. Further, the requirement that the generated captions are grammatically accurate and well-formed adds to the challenge of building intelligent image captioning systems. These systems, however, could greatly benefit those who are visually challenged to gain a superior sense of their neighborhood. With widespread access to mobile phones that are capable of taking photos on the go, the visually challenged user will be able to take a photo of his environment. A caption for the photo would be generated and read aloud to the user. In this paper, the focus is on the development of a deep recurrent neural network architecture-based model which can generate descriptive sentences that serve as captions for images. The model extracts the features from the image fed to it via a Convolutional Neural Network, which in turn is advanced to a Long Short-Term Memory network, that is an artificial recurrent architecture which produces a highly descriptive caption for the image fed in natural language. Training of the model focuses on achieving maximum similarity to a target sentence which is specified during the training phase; hence the model aims to achieve fluency in language and accuracy in identifying context and content of an image exclusively from the target descriptions provided for images in the dataset and use that to caption other images fed to it. The proposed model was evaluated quantitatively using BLEU scores and qualitatively as well. The model presented achieves high standards of accuracy and is capable of producing accurate, descriptive captions for images. The model, therefore, has immense potential to contribute to bettering the lives of those with visual impairments.

Read the paper · More papers on PaperTik