Visual-Semantic Dual Channel Network for Visual Question Answering
Xin Eric Wang, Qiaohong Chen, Ting Wei Hu, Qi Long Sun, Yubo Jia · 2021
Recently, the existing visual question answering (VQA) models based on the attention mechanism have achieved state-of-art results. However, attention-based networks only rely on question guidance to capture relevant image features, ignoring the high-level semantic information of the image. Therefore, the key challenge of the VQA task lies in obtaining effective semantic embedding and fine-grained visual understanding during the reasoning process. In this research, we propose a novel visual-semantic dual channel network to answer related questions from both visual and semantic perspectives. Specifically, the visual channel uses the relational reasoning method with an attention mechanism to capture visual objects and their relations, while the semantic channel can capture high-level semantic information from the global and local image by the semantic attention module. We confirmed the effectiveness of the proposed model and each module through extensive experiments on two versions of VQA datasets. Interpretability shows that the visual-semantic dual channel network can dynamically model to infer the most relevant answer to the question.