Zero-Shot Visual Question Answering based on DataSet Redistribution

Journal of System and Management Sciences · 2022

Visual Question Answering is an extremely active research area in which the computer is given an image, a question in natural language, and it is required to give a correct answer to the question according to the semantics of the input image.The ability of VQA system to answer new questions about unseen images during training process is one main measure of effectiveness of the VQA model and this capability is called Zero-Shot VQA, but VQA datasets suffer from some problems that hinder good evaluation on models trained on these datasets.Firstly, Testing instances are not chosen perfectly to address how much the trained model accomplish the task of asking about new concepts that is not presented during training process.Secondly, most of visual question answering datasets suffer from problems in their contents such as small dataset size, leakiness of explicitly defined question types, and question types have abused evaluation scores that makes it difficult to evaluate algorithms on them.So models are not perfectly evaluated on such datasets.In order to avoid those evaluation obstacles, experiment is done on TDIUC dataset which has explicitly defined 12 question types, data are redistributed for zero shot task by re-splitting it to new training, val, and test instances such that test instances contains new concepts that is not presented in training data.Evaluation is done using methods that give a more representative measure of accuracy over all question types( Simple Accuracy, AMPT, HMPT) and one more evaluation schema(GMPT) is proposed for evaluating accuracy which is more expressive.Experiment shows that evaluation results on TDIUC dataset before redistributing train, val, and test sets for Zero Shot purpose gives inaccurate indicator of model performance (around 20% higher performance)

Read the paper · More papers on PaperTik