Indic Visual Question Answering
Aditya Chandrasekar, Amey Vijay Shimpi, Dinesh Naik · 2022 IEEE International Conference on Signal Processing and Communications (SPCOM) · 2022
Visual Question Answering (VQA) is a problem at the intersection of Computer Vision (CV) and Natural Language Processing (NLP) which involves using natural language to respond to questions based on the context of images. The majority of existing methods focus on monolingual models, particularly those that only support English. This paper proposes a novel dataset alongside monolingual and multilingual models using the baseline and attention-based architectures with support for three Indic languages: Hindi, Kannada, and Tamil. We compare the performance of traditional (CNN + LSTM) approaches with current attention-based methods using the VQA v2 dataset. The proposed work achieves 51.618% accuracy for Hindi, 57.177% for Kannada, and 56.061% for the Tamil model.