Visual Question Answering Combining Multi-modal Feature Fusion and Multi-Attention Mechanism

Cai Linqin, Liao Zhongxu, Zhou Sitong, Chen Kejia · 2021

To enhance the semantic information and more accurately capture the image features in visual question answering (VQA) models, this paper presented a new VQA approach based on the multimodal features fusion and multiple level attention mechanism. Firstly, the ResNet network and Faster R-CNN were used to extract the local and global feature maps of the images; and then a two-layer LSTM was built to semantically encode text feature vectors. On this basis, Multi-modal Factorized Bilinear pooling approach was applied to fuse the image features and the text features. In addition, we combined the self-attention mechanism, the guided attention mechanism and the multi-head attention mechanism to form a multilayer attention network so as to more accurately obtain the required target features from images. The proposed models were tested and verified on the VQA-v2 data set. Experimental results show that the proposed methods are efficient and can obtain 67.18% of accuracy.

Read the paper · More papers on PaperTik