A Transformer-based Medical Visual Question Answering Model

Lei Liu, Xiangdong Su, Hui Guo, Daobin Zhu · 2022 26th International Conference on Pattern Recognition (ICPR) · 2022

While the Transformer architecture has been widely used in natural language processing tasks and computer vision tasks, its application in medical visual question answering is still limited. Most current methods rely on an image extractor to obtain visual features and a text extractor to capture semantic features, and then a fusion module to merge the information from the two modalities to predict the final result. In contrast, this paper proposes a novel Transformer-based medical vision question answering model, called MQAT, in which an improved Transformer structure is used for feature extraction and modal fusion to achieve better performance. Experimental results demonstrate that our Transformer structure not only ensures the stability of the model performance, but also accelerates its convergence, and the MQAT model outperforms the existing state-of-the-art methods.

Read the paper · More papers on PaperTik