Seeing With Sound: Automatic Image Captioning with Auditory Output for the Visually Impaired
P. Revathy, Sorna Shanthi D, M Bhavani, Priya Vijay · 2023
The human eye can discern imagery and add a mental annotation of it. However, the visually impaired community cannot visualize or label an image or environment. As an assistive technology for the visually impaired, the proposed system will automatically generate a highly relevant caption that accurately describes the events portrayed in an image. The generated caption will undergo further conversion and be presented as an auditory output for easy user accessibility. The captioning process is performed using Deep Learning algorithms – Convolutional Neural Network (CNN), InceptionV3 model, Recurrent Neural Network (RNN), and Long Short-Term Memory (LSTM) – and Beam Search.