SCAG: Semantic Co-occurring Attention Guided Alignment for Knowledge-based Visual Question Answering
Zheng Liu, Kunyu Yang, Yu Weng, Zheng He, Xuan Liu, Honghao Gao · ACM Transactions on Multimedia Computing Communications and Applications · 2025
In the realm of Knowledge-based Visual Question Answering (KB-VQA), the intricacy of the task lies in adeptly retrieving pertinent information from external sources and seamlessly aligning and amalgamating multimodal features. While numerous studies have effectively leveraged external knowledge to enrich factual connections among entities, there exists a tendency to overlook the significant reservoir of implicit information inherent in the visual-textual dimension. This oversight often results in suboptimal alignment and an undue reliance on the knowledge base. To address these challenges, this article introduces a novel strategy called SCAG. This approach aggregates the semantic co-occurring attention from diverse regions within images and various tokens within textual inputs using guidance weights to construct joint probabilistic representations grounded in the visual and textual dimensions, respectively. By employing this alignment strategy, the goal is to substantially mitigate information loss, reinforce inter-feature constraints within the model, reduce reliance on external knowledge sources, and enhance self-reasoning capabilities. The efficacy of our proposed model is comprehensively evaluated on the VQAv2 and OK-VQA datasets, with comparative analyses against multiple models conducted on the Ambiguous Knowledge (AK) dataset. Notably, our model exhibits a noteworthy 4.62% improvement over the state-of-the-art in addressing the knowledge dependency problem.