BPCN:A simple and efficient model for visual question answering

Feng Yan, Wushour Silamue, Yanbing Li, Yachang Chai · 2022

Visual question answering (VQA) is a cross modal task that combines computer vision tasks and natural language processing tasks. Due to the traditional attention mechanism model can not effectively solve the gap between the high-level semantics of words and the low-level abstract pixels of images, we use the question guided multi-head concatenated attention mechanism to map the question features and image features to the shared space, then obtain the question guided image attention and the image features related to the question. In addition, existing VQA models usually use static word vectors to encode text features. In this paper, based on the BERT's dynamic word vector encoding question, the accuracy of the model is further improved. Our model is simple and easy to train, with low complexity and significantly improved performance. Without using VG dataset for data enhancement, our model reached 69.28% on the test-std set.

Read the paper · More papers on PaperTik