Smart Assistant for Visually Impaired People using Deep Learning Algorithms
Parmar Bliss Bipin, S. Abirami · 2023
Vision impairment has a profound influence on quality of life. This proposed work aims to work as an eye for the blind people. It will take in pictures as an input and interpret the details in the given picture as captions. The challenge with this system is that it requires greater accuracy in extracting information from images and associating it with captions because captions for the blind must be clear, concise, and reliable. Deep learning teaches computers to perform human-like tasks through multi-layer architectures. To address the eyesight challenge for the blind population, this study proposes a hybrid deep learning autoencoder ensemble, which is a combination of Convolution Neural Network (CNN) with Gated recurrent units (GRUs) along with a novel attention mechanism to extract the inherent features of the image that are most closely related to the caption. Here, the CNN serves as the encoder and GRU serves as the decoder. The proposed attention mechanism uses dense layers to draws attention to areas of the image and grasp more relevant features of the image associated with the caption. Flickr8K dataset is used to assess the suggested model. which contains 8091 images, each of which includes five captions, for a total of 40455 captions. The efficacy of the proposed CNN-GRU model with attention layer to selectively attend to various components of the image while generating each word of the caption enables it to outperform the existing approach by 10.59% accuracy improvement by perceiving better Bilingual Evaluation Understudy (BLEU) score.