QASVQA: Overcoming Language Priors with Question Answer Semantics Guided Margin Loss
Haoyu Fang, Weidong Tian, Junjie Li · 2025
Visual Question Answering (VQA) models suffer from the language prior problem, which refers to blindly making predictions based on the language shortcut. Recent studies revisit this problem from the perspective of answer feature space by introducing margin loss to differentiate between answer feature spaces based on varying answer frequencies and answering difficulty. However, these methods fail to effectively allocate differentiated margins based on the semantics of individual samples, leading the model to rely predominantly on the language modality for answer reasoning. To address this issue, we propose a bias mitigation approach termed QASVQA (Question-Answer Semantic-guided Margin Loss for VQA), which effectively suppresses over-reliance on potentially biased samples. In addition, we introduce a Margin Adaptive Circle Loss, designed to impose varying degrees of penalties on intra-class and inter-class similarities during the separation of positive and negative sample features. This enhances the model's ability to discriminate between intra-class and inter-class differences. Experimental results on the VQA-CP v2, VQA v2, and VQA-CE datasets demonstrate the effectiveness of the proposed QASVQA approach.