Robust Knowledge Distillation and Self-Contrast Reasoning for Debiased Visual Question Answering

Ke Ning, Rongrong Shen, Zhixin Li · 2024

The language prior problem in VQA makes the model directly predictis based on questions, causing the model’s performance to drop sharply outside the distribution. Current debiased methods often achieve good out-of-distribution generalization capabilities at the expense of in-distribution performance degradation. We propose a novel method combining Robust knowledge distillation and self-contrast Reasoning (RR-VQA) to solve the language prior problem in VQA. We propose the QAS module to select reasonable questions for images, perform knowledge distillation through the ID and OOD teacher models, and obtain pseudo answers after passing through the QAS module. The CRSG module we propose synthesizes four visual language positive and negative samples for contrastive reasoning, which effectively increases the semantic dependency of the images while avoiding performance degradation due to spurious correlations between questions and answers. Our method is model-agnostic and achieves state-of-the-art performance on VQA-CP v2 dataset while maintaining performance on VQA v2 dataset.

Read the paper · More papers on PaperTik