Visual Question Answering: In-Depth Analysis of Methods, Datasets, and Emerging Trends
Pooja Singh, Munish Mehta, Gourav Bathla, Malathy Batumalay · 2025
Visual Question Answering (VQA) is a multi-modal technique that merges computer vision and natural language processing to enable systems to answer questions about images. Recently, several improvements have been made to VQA, driven by its wide applications. In this paper, we analyze and compare key studies, focusing on traditional and modern approaches to VQA. Traditional techniques mostly depend on Convolution Neural Networks (CNN), which extract visual features, while Recurrent Neural Networks (RNN) are used for text-based features. Further, in advanced models transformers, self-attention, and pre-trained CNN models are applied for better accuracy. In this paper, traditional and modern techniques are examined in detail for a better understanding of future researchers. Further, popular VQA datasets like VizWiz, COCO-QA, etc. are discussed for their strengths and limitations. In spite of improvements, VQA system faces issues like handling complex reasoning and achieving generalization among various domains. These open research problems and issues are presented in this paper.