Question Splitting and Unbalanced Multi-modal Pooling for VQA

Mengfei Li, Huan Shao, Yi Ji, Yang Yang, Chunping Liu · 2019

Visual question answering (VQA) is a cross-modal learning task that requires question understanding, image interpreting, and associating question with image. Most models generally did not consider to use different question parts in different modules, nor took into account the different roles of multi-modal features in fusion. In this paper, we proposed a question splitting and unbalanced multi-modal pooling approach. The question is split into two parts, one part contains object information called question footer. The other contains question type information called question header. Then we superimposed several layers of feature reinforcement linear overlaps on the basis of Multi-Modal Factorized Bilinear Pooling in order to give them different weights. Considering the interaction of multimodal features, our model also introduces the co-attention mechanism. Experimental results demonstrated our framework is superior to the previous models such as Oracle(GVQA, SAN) and QRU. The accuracy of our model increased from 61.96% to 64.44% on VQA 2.0 dataset and from 62.5% to 65.72% on COCO QA dataset.

Read the paper · More papers on PaperTik