Visual question answering based on spatial DCTHash dynamic parameter network
Changhong Liu, Mingwen Wang, Jihua Ye, Xiangshen MENG, Aiwen Jiang · Scientia Sinica Informationis · 2017
As research on deep learning and multimodal fusion continues to develop, question-answering systems have evolved from purely textual style to incorporating visual information. Visual question answering is becoming one of the inter-discipline topics between natural language processing and computer vision, and is receiving much attention. The dynamic parameter prediction network proposed by Hyeonwoo et al. is capable of effectively combining questions and visual information. However, when the network did weights hashing, the locations of the hashing codes were random, ignoring the image contents' spatial distributions. To overcome this shortcoming, in this paper, a new spatial DCTHash-based dynamic parameter prediction network on multimodal fusion is proposed for predicting visual question's answers. A conv7 feature mapping is extracted to retain the spatial visual information in a fully convolutional style. Then, question-related and visual structure-preserved convolution kernels are generated to perform the visual answer prediction. The proposed model has been compared with current commonly used algorithms on two public datasets: COCOqa and MSCOCO-VQA. The experimental results demonstrate that the proposed model has competitive advantages and can achieve relatively better performance.