Research On Visual Question Answering Based On Deep Stacked Attention Network
Zhu Xiaoqing, Junjun Han · Journal of Physics Conference Series · 2021
Abstract Aiming at the problem that the existing visual question answering model has a language bias in high-level logical reasoning, the model describes images or answers questions with low accuracy. A visual question answering model based on deep stacked attention networks is proposed. At the same time, the attention mechanism is introduced to help the model fully mine the image information; secondly, the problem attention mechanism is introduced to pay attention to the image and the question information at the same time when the problem feature is extracted; finally, the features of the image and the question are integrated and the answer is selected in the classifier. Taking the Visual Genome and VQA2.0 data sets as examples for empirical analysis, the results show that the accuracy of predicting answers is 1.13% higher than the existing best model, which proves the effectiveness and applicability of the model.