Speech Enabled Visual Question Answering using LSTM and CNN with Real Time Image Capturing for assisting the Visually Impaired
Annapurna P Patil, Amrita Behera, Palagati Anusha, Mitali Seth, Prabhuling · 2019
The proposed work benefits visually impaired individuals in identifying objects and visualizing scenarios around them independent of any external support. In such a situation, the user can capture a real time image of the surrounding and ask an open-ended question, classification question, counting question or yes/no question to the application by speech input. The proposed application uses Visual Question Answering (VQA) to integrate image processing and natural language processing which is also capable of speech to text translation and vice versa that helps to identify, recognize and thus obtain details of any particular image. The work uses a classical CNN-LSTM model where image features and language features are computed separately and combined at a later stage using image features and word embedding obtained from the question and runs a multilayer perceptron on the combined features to obtain the results. The model achieves an accuracy of 57 per cent. The model can also be utilized to develop cognitive interpretation better in kids. As the application is speech enabled it is best suited for the visually impaired with an easy to use GUI.