Multimodal Cross-guided Attention Networks for Visual Question Answering

Haibin Liu, Shengrong Gong, Yi Ji, Jianyu Yang, Tengfei Xing, Chunping Liu · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2018

Visual Question Answering (VQA) is an attractive topic combining computer vision with natural language processing.It is more challenging than text-based question answering because of its multimodal nature.The VQA reasoning process requires both effective semantic embedding and fine-grained visual comprehension.Existing approaches predominantly infer answers from visual spatial information, while neglecting important semantic information in questions and the guidance information between images and questions.To remedy this, we imitate the human mechanism of cross-reasoning about visual and textual information and propose a multimodal cross-guided attention network (MCAN) for VQA which employs a cross-guided joint learning strategy with a gated activation learning method, which can simultaneously capture both rich visual spatial information and significant semantic information.We evaluate the proposed model on two public datasets: VQA dataset and COCO-QA dataset.Extensive experiments show state-of-the-art performance on the datasets.

Read the paper · More papers on PaperTik