Adapted GooLeNet for Visual Question Answering
Huang J. Jie, Yue Hu, Weilong Yang · 2018 3rd International Conference on Mechanical, Control and Computer Engineering (ICMCCE) · 2018
Visual Question Answering (VQA) aims at answering a question about an image. In this work, we introduce an effective architecture –Adapted GooLeNet (AG)– into a typical VQA method MUTAN instead of LSTM for question features capturing. This improvement can capture more levels of language granularities in parallel, because of the various sizes of filters in AG. The empirical study on the benchmark dataset of VQA demonstrates that capturing sentence features on different levels of granularities benefit sentence modelling by utilizing AG.