Visual Question Answering: Attention Mechanism, Datasets, and Future Challenges
Minyi Tang, Ruiyang Wang, Siyu Lu, Ahmed Alsanad, Salman Ali AlQahtani, Wenfeng Zheng · 2024
Visual Question Answering (VQA) is a multidisciplinary research problem that integrates computer vision and natural language processing, thereby advancing the boundaries between these fields. Recently, VQA has garnered significant attention from researchers in both computer vision and natural language processing domains. In a VQA system, an image and a related question in natural language serve as inputs, with the system required to provide an output that is both accurate and based on natural language processing. This paper reviews the concepts, current state, and improvement methods of VQA research, with a focus on aspects such as attention mechanisms, datasets, and the primary challenges facing VQA. Additionally, bibliometric methods are employed to analyze the VQA field. The article also explores potential future research directions. The adaptation of VQA to dynamic images and videos for daily applications is anticipated to hold significant developmental potential. Furthermore, to address complex questions, more fine-grained modal models are expected to be developed, and systems capable of providing multiple answers in VQA are likely to become increasingly prevalent.