Language-guided Bias Generation Contrastive Strategy for Visual Question Answering

Enyuan Zhao, Ning Xi Song, Ze Zhang, Jie Nie, Xinyue Liang, Zhiqiang Wei · ACM Transactions on Multimedia Computing Communications and Applications · 2025

Visual question answering (VQA) is a challenging task that requires models to understand both visual and linguistic inputs and produce accurate answers. However, VQA models often exploit biases in datasets to make predictions rather than reasoning based on the inputs. Prior approaches to debiasing have suggested the implementation of a supplementary model, deliberately designed to exhibit bias, which subsequently informs the training of a resilient target model. Nevertheless, such techniques merely quantify the model’s divergence based on the statistical distribution of labels within the training dataset or in relation to unimodal branches. In this work, we propose a novel method of generating bias from the target model itself, called LEGO, which aims to combat the language guidance bias. Specifically, LEGO framework employs a generative network that assimilates the biases inherent in the target model by integrating adversarial goals with the principles of knowledge distillation. Then, we use a debiased contrastive learning strategy to model the language guidance bias of caption and question. In the process of modeling, in order to obtain robust semantic coreference, the multimodal representations of two semantic granularity are modeled by mutual information fusion and contrast learning difference modeling. We evaluate our method on various VQA-biased datasets, including VQA-CP2, GQA-OOD, and RSICD, and show that it outperforms similar methods.

Read the paper · More papers on PaperTik