Multi-view Attention Networks for Visual Question Answering
Min Li, Zongwen Bai, Jie Deng · 2024
Visual question answering (VQA) is a typical multimodal task that necessitates a combination of computer vision and natural language processing expertise. The fundamental essence of VQA lies in the simultaneous comprehension of fine-grained language and visual information. In recent years, transformer-based methods have exhibited remarkable success in advancing the state-of-the-art in VQA. In this paper, we present an enhanced model, namely the Multi-view Attention Network (MVAN), a variant of the Transformer architecture. MVAN improves the effect of the model in filtering irrelevant information and focusing on local features. Specifically, we augment our network with a Gated Linear Unit (GLU) to discern and filter irrelevant or inconsequential information. Additionally, a Gated Convolution Block (GCB) is introduced to the self-attention layer of the Transformer variant. This integration facilitates the extraction of contextual semantic image information by considering both channel and spatial perspectives. As a result, the model effectively combines both local and global information, hence improving its predicted accuracy in VQA tasks. Ultimately, the model is subjected to verification and testing using the VQAv2 dataset. The outcomes of these evaluations demonstrate a notable enhancement in the performance of our model when compared to the existing methods. Furthermore, we also conduct extensive ablation experiments to explore the reasons for the effectiveness of the MVAN.