A Co-Attention Mechanism for Medical Visual Question Answering on the VQA-Med 2019 Dataset

Omar.M Kasba, Fady Sameh Farahat, Moustafa Karam Abd El Mohsen, Ahmed Mohamed Sofi, Mohamed Qadri Salah, Nehal Khaled · 2024

This paper introduces a novel AI-based architecture for Medical Visual Question Answering (VQA). Our approach leverages advanced visual and textual feature extraction techniques, integrating them using a unique Multi-modal Factorized Bilinear (MFB) pooling mechanism. Evaluated on the VQA-Med 2019 dataset, the proposed model achieved an overall classification accuracy of 0.639. The experimental results demonstrated that the proposed method has superior performance compared to existing methods on the VQA-Med 2019 dataset.

Read the paper · More papers on PaperTik