Medical Visual Question Answering using Contrastive Language-Image Pre-training
Zahra Khalid Butt, Aiman Shahid Butt, Tehseen Zia, Mutahra Khalid Butt · 2024
Medical Visual Question Answering (MedVQA) is an emerging field that combines natural language processing and computer vision to enable systems to answer questions about medical images. Despite its potential to enhance diagnostic accuracy and support medical professionals, MedVQA faces significant challenges. Two major challenges are the lack of annotated data and the need for AI explainability. Existing models often treat MedVQA as an answer classification task, which restricts their capacity to adapt to a variety of questions and real-world situations when predefined answers are unavailable. In this paper, we proposed a novel generative MedVQA model that converts visual information into textual data using the CLIP model, improved by our suggested Contrastive Concept-Phrase Pre-training (C2P2) technique. This textual information together with the query is then fed into a language model, which produces human-like responses. We address interpretability by utilizing LIME to highlight significant image regions, which improves transparency. In addition, we create a dataset from the IU medical report collection to enable more complex, natural language responses. We compared our model with state of the art models, and our model performed better than the existing models. Our work provides a substantial contribution to increasing the adaptability, interpretability, and practical implementation of MedVQA systems in healthcare.