Computer Vision and Voice Assisted Image Captioning Framework for Visually Impaired Individuals using Deep Learning Approach
K M Safiya, R. Pandian · 2023
Blind or visually challenged have considerable obstacles when obtaining visual material, restricting their capacity to derive meaningful information from pictures. This research proposing a novel framework combining computer vision with voice-based image captioning using deep learning techniques. The proposed system employs VGGNet-16 model for photo processing and long short-term memory networks (LSTMs) for natural language processing. The models used in this research were trained on Flickr8k, Flickr30k, and a bespoke dataset. Subsequently, these models were deployed on a Raspberry Pi 4B single-board computer that was equipped with a graphics processing unit. A two-fold approach is used, whereby the first phase entails exposing the input image to preprocessing using a pre-established VGGNet-16 model. This methodology enables the retrieval of relevant visual attributes, therefore capturing the intrinsic semantic content of the picture. The aforementioned attributes are fed into a language model based on LSTM to generate explanatory captions. To enhance communication with visually impaired people, the generated captions are converted into audible speech using text-to-speech synthesis technology. The effectiveness of the proposed architecture is evaluated via extensive experiments using benchmark datasets and real-time images obtained with a NoIR camera unit. The generated captions are evaluated using quantitative assessment criteria, namely BLEU and ROUGE scores. The developed VGGNet16 model has exceptional performance in terms of accuracy (95.620%), precision (96.928%), recall (87.879%), and F1 score (92.182%). The results suggest that using computer vision in conjunction with voice-based image captioning offers a promising solution for addressing the challenges faced by those with visual impairments in perceiving and comprehending visual content.