Voice driven Visual Question Answering
Abhijith Santhosh, K Rohith, Ananthakrishnan S, A Thakker Rajesh, Jasmine Bhaskar · 2024
Visual Question Answering is a multi-modal task to find an answer to any question asked about a given image. In this paper, we propose a model of dual CNNs that provide the correct answer to a spoken question based on a given image in audio or text form. Mirroring real-world scenarios like helping the visually impaired, the questions and answers are considered in spoken mode. Visual questions selectively address different areas of an image, including underlying context and background information. In this paper, we introduce a dual-CNN model—a conjunction of two convolutional neural networks concerning Visual Question Answering. This first CNN model is used to select the significant features from the image that contribute to the Answer, and the second CNN to extract relevant textual features from the question. The proposed framework has been tested with VQAV2.dataset. In terms of overall accuracy and execution time, the experimental findings demonstrated very good results.