BERT based Multiple Parallel Co-attention Model for Visual Question Answering
Mario Dias, Hansie Aloj, Nijo Ninan, Dipali Koshti · 2022 6th International Conference on Intelligent Computing and Control Systems (ICICCS) · 2022
Humans can easily interpret visual and textual content whereas for a machine this is a challenging task. Visual question answering is a well-known problem in the field of computer vision and NLP where an image and a question related to the image are given and the machine has to generate a natural language answer. This paper explores the use of transformer BERT as a language model in VQA. The BERT-based multiple parallel co-attention visual question answering model has been proposed and the effect of introducing a powerful feature extractor like BERT for language modeling has been studied. From the experimental results, it is concluded that the proposed model improves over the original baseline VQA model by 3%.