Voice Assisted Image Captioning and VQA For Visually Challenged Individuals

Kristen Pereira, Rushil Patel, Joy Almeida, Rupali Sawant · 2022 IEEE 19th India Council International Conference (INDICON) · 2022

Vision is one of the most vital senses for a person’s well being. More than 285 million people suffer from the problem of visual impairment. This count is predicted to be three fold in coming 30 years. Independent navigation and safe travel are difficult for these individuals, as they have difficulty perceiving information from their surrounding and communicating. Rather than relying on visual cues to guide blind people, the proposed project will help them navigate the world through the use of audio means. It will empower visually impaired individuals to explore independently by using the system to detect objects in their vicinity and without any outside assistance. Within our proposed research, we have built a mobile application which leverages the power of image processing, and deep learning techniques to identify and describe the current scene through the camera and inform it to the user by audio cues. As a result of not being able to distinguish between different objects, the already existing approaches have been limited, resulting in low performance and accuracy. With this we attempt to provide enhanced performance, better accuracy and hence a more practicable and reliable alternative.

Read the paper · More papers on PaperTik