Transformer Gate Attention Model: An Improved Attention Model for Visual Question Answering
Haotian Zhang, Wei Wu · 2022 International Joint Conference on Neural Networks (IJCNN) · 2022
Given an image and an opened-ended question related to the picture, the Visual Question Answering (VQA) model aims to provide the right answer to the question for the image. This is a challenging task that requires a fine-grained simultaneous realization of the visual content of the image and the textual content of the question. However, most current models ignore the region-to-word interaction and the noise effect of irrelevant words. This paper proposes a novel model called Transformer Gate Attention Model (TGAM) to capture the inter-modal information dependence to solve the above problems. Specifically, TGAM is composed of Adaptive Gate (AG) and Parallel Transformer Module (PTM). AG fuses information between different modalities and reduces noise, and PTM obtains more advanced cross-modal feature representation. Many qualitative and quantitative experiments were conducted on VQA-v2 dataset to verify the validity of TGAM. And extensive ablation studies have been conducted to explore the reasons behind the effectiveness of TGAM. Experimental results show that the performance of TGAM is significantly better than previous state-of-the-art technologies. Our best model achieved 71.28% overall accuracy on the test-dev set and 71.6% on the test-std set.