Adapted GooLeNet for Visual Question Answering

Huang J. Jie, Yue Hu, Weilong Yang · 2018 3rd International Conference on Mechanical, Control and Computer Engineering (ICMCCE) · 2018

Visual Question Answering (VQA) aims at answering a question about an image. In this work, we introduce an effective architecture –Adapted GooLeNet (AG)– into a typical VQA method MUTAN instead of LSTM for question features capturing. This improvement can capture more levels of language granularities in parallel, because of the various sizes of filters in AG. The empirical study on the benchmark dataset of VQA demonstrates that capturing sentence features on different levels of granularities benefit sentence modelling by utilizing AG.

Read the paper · More papers on PaperTik