Voice Enabled Deep Learning Based Image Captioning Solution for Guided Navigation

Senthil Kumar A, Selvaraj Kesavan, J. Earnest Jayakumar, Ananda Kumar K S, Prasad Maddula · 2023

The use of technology to assist visually impaired individuals is crucial in addressing the global issue of vision impairment. Worldwide more than billion people suffer from a vision impairment that should have been avoided or is yet unaddressed. According to the statistics, there is a significant need for solutions that can help those who are visually impaired, mainly in the middle- and low-income countries where the vision impairment population is higher. It is anticipated that population expansion and ageing will increase the likelihood that more people may get vision impairment. The efficientnetB3 deep learning algorithm will be used in this project to caption images for blind people. so, they can learn about object identification, distance, and position. This has been accomplished by utilizing advanced picture captioning techniques, efficient net B3 algorithms, and tokenization approaches, where the computer learns the scenes with various captions. The computer recognizes and forecasts any image that is acquired using the camera. The significant objects are also anticipated, and the camera's distances are determined. Following the prediction, the user receives an audio output that can be used to determine the object's position and distance. Hence, with the aid of this research, we give the blind artificial eyesight that can give them confidence when they move on their own. The aim is to step forward in addressing the global issue of vision impairment. The use of technology to assist visually impaired individuals is crucial in providing them with the tools they need to navigate their environment and live their lives with greater ease. By utilizing advanced algorithms and image captioning techniques, the quality of life can be improved for people worldwide who are affected by vision impairment. The intension is to develop an artificial vision for vision impaired people by detecting real time objects, distance and the position of it from the person using Audio Output and to develop a model for image captioning to predict the captions.

Read the paper · More papers on PaperTik