Overcoming Data Limitations and Cross-Modal Interaction Challenges in Medical Visual Question Answering

Yurun Bi, Xingang Wang, Qingao Wang, Jinxiao Yang · 2024

Medical Visual Question Answering (Med-VQA) is a sophisticated multimodal technology within the medical domain that utilizes natural language processing and artificial intelligence techniques. It serves as a tool in the medical field to provide information and answers, aiming to correctly respond to medical questions related to specific professional medical images. The primary goal is to assist patients in understanding their physical condition and provide decision support for healthcare professionals. While Generalized Visual Question Answering has seen significant advancements, challenges persist in the medical domain. Firstly, medical data has inherent limitations, such as the scarcity of annotated medical datasets and the complexity of medical images (including factors like noise, artifacts, and low contrast). Secondly, in terms of cross-modal fusion, there exists an intricate relationship between medical images and textual information. The complexity of the relationship between medical images and medical text, as well as the incorporation of global contextual details, pose challenges. To address these issues, this paper proposes the Feature-Enhanced Network with Conditional Mixing. This innovative approach aims to overcome the scarcity and complexity of medical images without the need for additional datasets. Simultaneously, to tackle the challenge of cross-modal fusion and information interaction and to address the complex relationship between medical images and text, we propose the Cross-Adaptive Fusion Network. Experimental results conducted on the VQA-RAD dataset demonstrate noticeable improvements in accuracy compared to mainstream Med-VQA models.

Read the paper · More papers on PaperTik