Cascading Top-Down Attention for Visual Question Answering
Weidong Tian, Rencai Zhou, Zhong‐Qiu Zhao · 2020
For solving Visual Question Answering (VQA), we commonly employ images and questions simultaneously to predict answers. Some attention mechanisms should be used to focus on the most valuable information, because there are too much information extracted from images and questions. Top-Down Attention (TDA) is one of the famous attention mechanisms. For standard TDA, only important regions of the image associated with the question are highlighted. In this work, we propose a Cascading Top-Down Attention (CTDA) model. CTDA highlights the most important information collected from images and questions by a cascading attention process. First, the key words of the question, associated with the image, are highlighted by using a Question Top-Down Attention (QTDA). Then, important regions of the image, associated with the question, are highlighted by using of Image Top-Down Attention (ITDA), useless information of the images and questions can be ignored effectively. We evaluate our model on two popular VQA data sets. CTDA obtains better results than standard TDA and the other state of the art models.