A Review on VQA: Methods, Tools and Datasets

Mayank Agrawal, Anand Singh Jalal, Himanshu Sharma · 2023

An new area called “visual question answering” (VQA) seeks to integrate CV with NLP. In order to get correct results, it entails creating models that can comprehend both textual questions and visual input, represented by videos or images. Applications for VQA systems include content-based image retrieval, medical image analysis, autonomous vehicles, and human-computer interaction. However, there are a number of difficulties in developing efficient VQA models, such as ambiguity in questions, sophisticated reasoning, processing multi-modal data, and data bias. To solve these issues and enhance the functionality and interpretability of VQA models, researchers are consistently investigating novel techniques, such as attention mechanisms, fusion tactics, and transformer-based architectures. In addition to a review of the existing problems and potential future developments in the area of visual question answering, this paper offers an overview of VQA methods, datasets, and tools. Finally, we go through potential routes for VQA and image comprehension research in the future.

Read the paper · More papers on PaperTik