An Improved Medical Visual Question Answering Model Based on CLIP and BERT
Umair Javed, Touqeer Abbas, Muneeb Raza, Faisal Mehmood, Jahanzaib Yaqoob, Hui Li · 2024
Visual Question Answering (VQA) is a versatile tool applicable in autonomous vehicles for manufacturing use cases up to recommendation platforms for e-commerce, thus greatly increasing process efficiency, sophistication of decisions, and user engagement. One area that stands out is healthcare, where this could revolutionize diagnostic procedures and patient comprehension, which will help patients at large readings more effectively. Reduced convergence time and enhanced interpretability because solution paths are directly enforced through constraints. Coherent question answers are difficult to form, yet, remain a key challenge. In this work, we present a framework for knowledge-based VQA utilizing answer heuristics. Then, two types of answer heuristics are presented: answer candidates and answer-aware examples. Answer candidates and also answer-aware examples with the same answer as the test input, which was a downside as well as a limitation in our benchmark were currently provided which were aimed to make our experiment extra straightforward. Our proposed MedVQA version with CLIP-based visual attributes along with BERT to enhance the MCAN model was able to increase the model to an accuracy of 78.89% (11.09+) over the current state-of-the-art. This is a significant development, proving that state-of-the-art models and visualization elements lead to improved system performance. The original framework intended to facilitate adoption and thereby improve healthcare decision-making and patient outcomes.