Cross-Guided Attention and Combined Loss for Optimizing Gastrointestinal Visual Question Answering
Qian Yang, Lijun Liu, Xiaobing Yang, Li Liu, Wei Peng · 2024
Medical Visual Question Answering (Med-VQA) aims at answering clinical questions related to medical imaging, providing important reference for medical imaging diagnosis. Gastrointestinal visual question answering can deeply understand and analyze gastrointestinal medical images, provides supporting information to answer questions about the diagnostic process, and improves diagnostic accuracy and efficiency. However, most existing Med-VQA can only handle simple and general questions, struggling to tackle visual question answering related to complex diseases. Therefore, this paper proposes a Med-VQA model called CACL, which uses cross-guidance and improved combined loss. The model focuses on gastrointestinal image analysis and achieves efficient and accurate feature extraction for gastrointestinal images. A cross-guided attention module is also designed to enhance the model’s reasoning ability when dealing with complex cross-modal tasks. In addition, a multi-task composite loss function is designed to balance the loss of segmentation task and classification task, and improve the overall performance. Experimental results show that the proposed method can effectively improve the accuracy of gastrointestinal visual question answering.