Cross-Modal Noise Information Elimination Network for Medical Visual Question Answering

Wanjuan Yuan, Chen Qiu · 2024

Medical Visual Question Answering (Med-VQA) aims to answer clinical questions given medical images and deduce the best answer. The existing Med-VQA models fuse question and image features by the dense interactions between each image region and each question. However, the existing Med-VQA methods ignore the noise information generated by the interactions between the image regions unrelated and the question words unrelate, but these noise information interfere with the inference of the model. We propose a Cross-Modal Noise Information Elimination Network (CMNIEN) for medical visual question answering. Firstly, question features are extracted by the BERT model, and image features are extracted by the ResNet network; Then, the question and image features were concatenated together, and a threshold based feature fusion module for Med-VQA was designed to eliminate the noise information generated by the interaction of different modal features of images and questions; Finally, a cross modal self attention module is used to enhance the attention of key features. Extensive experiments were conducted on the Med-VQA-2019 and VQA-RAD datasets, and the results showed that each improvement can effectively enhance model performance. When the threshold is set to 0.1, the overall accuracy of the model on the Med-VQA-2019 dataset reaches 81.9%, and the overall accuracy on the VQA-RAD dataset reaches 77.2%.

Read the paper · More papers on PaperTik