Real-Time Photo Captioning for Assisting Blind and Visually Impaired People Using LSTM Framework

K M Safiya, R. Pandian · IEEE Sensors Letters · 2023

This letter aims to tackle the problem of the visually impaired community by introducing an innovative framework that integrates computer vision with voice-based photo captioning via deep learning methodologies. In this research, a performance of three models was conducted, namely VGGNet-16, ResNet-152, and MobileNet-V3, to identify the best-performing model for hardware deployment. The models are trained using the Flickr30k, Flickr8k, and 2k custom datasets. Subsequently, the best-performed VGGNet-16 model is deployed on a Raspberry Pi 4B single-board computer with a graphics processing unit. To facilitate communication with visually impaired people, the generated captions are transformed into audible speech via text-to-speech synthesis technology. The VGGNet-16 model has superior performance compared to the other two models. The quality of generated captions is assessed using quantitative assessment criteria such as BLEU, BLEURT, and ROUGE scores. Also, input from the visually impaired to assess the usefulness, accessibility, and influence of the framework.

Read the paper · More papers on PaperTik