Multi-Granularity Cross-Attention Network for Visual Question Answering

Yue Wang, Wei Gao, Xinzhou Cheng, Xin Wang, Huiying Zhao, Zhipu Xie, Lexi Xu · 2023

Visual Question Answering (VQA) is a recent hot topic that involves multimedia analysis, computer vision (CV), natural language processing (NLP), and even a broad perspective of artificial intelligence, which is challenging and has obtained increasing attention. VQA needs a complete understanding of the spatial relationship, textual clues, as well as the common sense for an actual image. However, most existing approaches simply embed and concatenate the features of questions and images to predict answers. Treating all embeddings equally without consideration of relation consistency hinders the model performance. In this paper, we propose an explicit Multi-Granularity Cross-Attention network (MGCAN) that mutually learns the multi-modal branches. MGCAN jointly matches word-level representation with whole image, and patch-level representation with the whole question that infers the high-order vision-semantic relationship. Experiments conducted on VQA datasets demonstrate that the proposed MGCAN outperforms previous baselines. The cross-attention mechanism explicitly exploits the relevant visual and textual clues that lead to superior prediction.

Read the paper · More papers on PaperTik