Towards Precision Healthcare: Leveraging Pre-trained Models and Attention Mechanisms in Medical Visual Question Answering for Clinical Decision Support

Arvind Kumar, Anupam Agrawal · 2024

In the current landscape of healthcare technology, Medical Visual Question Answering (Med-VQA) holds significant promise for enhancing clinical decision-making through the integration of natural language processing and computer vision. Our research presents a novel approach leveraging BLIP models to advance Med-VQA capabilities. This unified framework seamlessly combines textual and visual information, enabling precise responses to complex medical queries. By employing BLIP processors for text and image encoding, we effectively integrate textual queries with relevant visual data from the PathVQA dataset. Our BLIP-based QA model is fine-tuned using the AdamW optimizer with a learning rate of 5e-5, ensuring efficient convergence. The model incorporates advanced Attention Mechanisms, as well as Coarse and Fine Attention modules, to enhance feature fusion and optimize prediction accuracy. Through rigorous evaluation on both training and validation datasets, our approach demonstrates competitive performance metrics, showcasing its ability to generate accurate and contextually relevant answers across diverse medical image inputs. The results of our comparative analysis reveal that our BLIP-based model achieves superior accuracy and robust performance compared to existing methods. This underscores the potential of BLIP models in significantly advancing Med-VQA research and applications, ultimately contributing to improved healthcare outcomes and more informed clinical decision-making. Our work highlights the efficacy of leveraging cutting-edge AI technologies to address complex challenges in the medical field.

Read the paper · More papers on PaperTik