Hypergraph-Based Model for Visual Question Answering with External Knowledge Integration

Chaoxing Li, Yuanqing Li, Yinghua Li, Yanan Zhang, Dianwei Wang · 2025

Visual Question Answering (VQA) serves as a bridge between computer vision and natural language processing, aiming to enable machines to achieve human-level understanding when observing images and reading text. To advance more complex cross-modal understanding and reasoning processes, most existing works primarily focus on integrating and clarifying multimodal knowledge, yet rarely address the deeper interactions and intricate relationships between multimodal knowledge. To address this issue, a Hypergraph Knowledge Enhanced Network (HKEN) model is proposed in this paper. By leveraging hyperedges for matching and updating node representations, HKEN tackles the limitations in understanding the inherent high-order semantics and reasoning across multi-hop relationships in external knowledge base. Innovatively, the introduction of a Transformerbased hypergraph attention mechanism contributes to effectiveness in handling complex visual question answering tasks, especially when the answer requires external knowledge or deductive reasoning. The experimental results on KVQA and FVQA data sets are 88.23% and 77.02% respectively, which are higher than the existing models.

Read the paper · More papers on PaperTik