Regulating Balance Degree for More Reasonable Visual Question Answering Benchmark

Ken Lin, Aihua Mao, Jiangfeng Liu · 2022 International Joint Conference on Neural Networks (IJCNN) · 2022

Superficial linguistic correlations is a critical issue for Visual Question Answering (VQA), where models can achieve high performance by exploiting the connection between question and answer, but fail to obtain better generalization ability for out-of-domain data. To ease such issue, VQA-CP v2.0 greedily re-partitions the distribution of VQA v2.0's training and test divides, it suppresses the performance improvement acquired by superficial linguistic correlations. However, some opportunistic methods (such as inverse supervision) can take advantage of the dataset's distribution characteristics to obtain high performance, which is incompatible with academic efforts to increase the model's visual reasoning and modal fusion abilities. To address this problem, we propose a more reasonable dataset in which we attempt to make the training split conform to the long-tailed distribution and the test split more balanced, so that inverse supervision does not result in performance gains and superficial linguistic correlations still can not assist the model in achieving high accuracy. Besides, we propose a decoupled training schema which can obtain better representation and visual reasoning modules to compensate for the shortcomings of ensemble-based methods that selectively learn some samples. Without any further annotations, such schema achieves state-of-the-art performance. In VQA-CP v2.0, it outperforms the simple baseline model UpDn by 15.54%. And its accuracy on VQA v2.0 has almost no drop compared to UpDn. Code is available at https://github.com/asklvd/new-benchmark-for-robust-VQA.

Read the paper · More papers on PaperTik