A Multimodal Vision-Language Framework for AI-Driven Medical Question Answering System

Kumkum Biswal, Saravanapriya Manoharan · 2025

The integration of imaging in medicine with VQA technology posed a significant challenge due to accuracy and diagnostic precision issues. This, however, has been further simplified with the introduction of AI-injected diagnostic techniques, which have enhanced everyday steps for medical practitioners. To tackle these issues, we suggest a model that utilizes cutting-edge deep learning frameworks to enhance question answering based on medical images. The model presented in this paper incorporates Vision Transformer (ViT-32), Bootstrapped Language-Image Pretraining (BLIP), and Multi-Scale Attention (MSA) feature encoders, which work together through an efficient multimodal fusion approach using cross-attention and concatenation. Leveraging deep learning architectures, the proposed model enhances medical image-based question answering by extracting, fusing, and interpreting multimodal features, enabling accurate and context-aware diagnostic responses. Our model shows an enhancement in accuracy by 4 to 6 percent, and an average augmentation in the F1 score of over 5 percent compared to the existing models. These progressions improves the accuracy of diagnosis, lowers ambiguous diagnosis, and improves the support for clinical decision-making. These findings mark a significant improvement over state-of-the-art models, providing better reliability in diagnosing errors, minimizing interpretation errors, and advancing AI-assisted medical VQA.

Read the paper · More papers on PaperTik