Adaptive Fusion for Visual Question Answering: Integrating Multi-Label Classification and Similarity Matching
Zhengtao Yu, Jia Zhao, Huiling Wang, Chenliang Guo, Tong Zhou, Chongxiang Sun · 2023
Visual Question Answering (VQA) is an important multimodal task in which models are required to answer questions based on visual cues. However, most visual question-answering models suffer from the language prior problem, which is caused by data bias. Specifically, VQA models tend to output high-frequency answers to answer questions while ignoring the information contained in the images. Many approaches have emerged to solve the language prior problem. However, previous approaches could only improve the performance of easy classes, and there is no means to solve hard classes effectively. In this paper, we will utilize more semantic information to guide the model for learning and better handle hard questions. Specifically, in addition to the classification task, we map the image question pairs and the answers to the same dimensional space and construct a similarity metric between the two to get the answers’ similarity-matching output. Moreover, we learn a set of parameters to fuse the classification output with the answers’ similarity-matching output, and finally, we use the fused output for prediction. We use answer weighting for each output to mitigate the language priors in computing the loss function. Moreover, we use answer masks for the classification outputs. Experimental results demonstrate the effectiveness of our method, which achieves a state-of-the-art performance of 62.20% on VQA-CP v2.