Augmenting Visual and Linguistic Understanding for MedVQA
Yongpei Ma, Jinman Kim · 2024
Medical visual question answering (MedVQA) is a challenging task that requires both visual and linguistic understanding of medical images and their questions. However, existing works generally rely on basic, simple language models that cannot interpret medical terms well and are limited in their ability to express the same linguistic meaning using grammar variations and word order. In this study, we propose a new MedVQA data augmentation framework comprising two novel methods: (1) MAA (MedVQA data augmentation AI-agent) - a data augmentation module that uses an AI-agent to transform the question format from interrogative to declarative and thus making it closer to the pre-training data of contrastive language-image pretraining (CLIP) models; and, (2) SNN-CLIP, a new deep learning architecture that exploits the advantages of temporal encoding in spiking neural networks (SNNs) to enhance the CLIP model for VQA. We conducted extensive experiments on MedVQA benchmark datasets and demonstrated that combining MAA and SNN-CLIP can effectively augment the original data, outperforming the state-of-the-art MedVQA as well as alternative data augmentation methods.