Leveraging Convolutional Models as Backbone for Medical Visual Question Answering
Lei Liu, Xiangdong Su, Guanglai Gao · 2024
Convolutional neural networks (CNNs) have made significant contributions to computer vision and offer the advantages of higher training efficiency and lower model complexity. However, their application as the backbone in medical visual question answering (MedVQA) remains an open question. To address this issue, we employ popular convolutional models, including ResNet, DenseNet, and ShuffleNet, as the foundation for MedVQA, achieving outstanding performance. Different backbones can be tailored to diverse real-world scenarios. The central challenge in utilizing CNNs for visual question answering is effectively managing textual features and integrating multi-modal information. To overcome this challenge, we design a novel global interaction attention (GIA) that facilitates efficient interactions between text and image features. Additionally, we utilize the dot product before the classifier output to enhance visual and textual modal fusions. To further enhance model performance, we propose a novel multi-modal hidden mixup (MHidMix) technique for data augmentation, which involves interpolating hidden states during model training. This data augmentation technique smoothes the decision boundary without the need for complex sample selection, further improving model performance. Experimental results underscore the versatility of our proposed framework across various convolutional models, leading to outstanding performance on four MedVQA datasets. Notably, we achieved an accuracy increase of 9.4% on the PathVQA dataset and 4.5% on the OVQA dataset.