An Visual Question Answering System using Cross-Attention Network
R. Dhana Lakshmi, S. Abirami · 2023
In this paper, we present a novel approach for tackling the challenging task of Visual Question Answering (VQA). Our method combines sparse cross-attention and masked self-attention techniques to improve the accuracy of answer prediction by effectively integrating visual and textual information. By leveraging cross-attention mechanisms, we establish strong relationships between visual and textual features, enhancing the model’s ability to comprehend the question and the visual content. Additionally, we introduce masked self-attention to further refine the answer prediction process. This module selectively focuses on relevant textual features while disregarding the answer, enabling the model to concentrate on the most informative aspects of the question. To evaluate the effectiveness of our proposed approach, we conduct experiments on the COCO VQA dataset. We compare our method against state-of-the-art approaches and demonstrate its competitive performance in terms of accuracy. Overall, our approach offers a promising solution for improving the efficiency, interpretability, and accuracy of VQA systems.