AIML at VQA-Med 2020: Knowledge inference via a skeleton-based sentence mapping approach for medical domain visual question answering

Zhibin Liao, Qi Wu, Chunhua Shen, Anton van den Hengel, Johan W. H. Verjans · Adelaide Research & Scholarship (AR&S) (University of Adelaide) · 2020

In this paper, we describe our contribution to the 2020 Im-ageCLEF Medical Domain Visual Question Answering (VQA-Med) challenge.Our submissions scored first place on the VQA challenge leaderboard, and also the first place on the associated Visual Question Generation (VQG) challenge leaderboard.Our VQA approach was developed using a knowledge inference methodology called Skeleton-based Sentence Mapping (SSM).Using all the questions and answers, we derived a set of classifiable tasks and inferred the corresponding labels.As a result, we were able to transform the VQA task into a multi-task image classification problem which allowed us to focus on the image modelling aspect.We further propose a class-wise and task-wise normalization facilitating optimization of multiple tasks in a single network.This enabled us to apply a multi-scale and multi-architecture ensemble strategy for robust prediction.Lastly, we positioned the VQG task as a transfer learning problem using the VGA task trained models.The VQG task was also solved using classification.

Read the paper · More papers on PaperTik