Visual Question Answering based on multimodal triplet knowledge accumuation
Fengjuan Wang, Gaoyun An · 2022
Knowledge-based visual Question Answer needs to be associated with external knowledge to better understand. Existing solutions generally acquire relevant knowledge from pure text knowledge bases, but these knowledge bases only contain some facts for simple questions and answers, lacking visual content for deep understanding. This paper proposes an explicit triplet to represent multimodal knowledge, which associates visual objects and factual answers with implicit relationships. The header is obtained by calculating the Euclidean distance between the regional feature and the text token. To narrow the heterogeneity gap, we propose three target losses of learning triplet representation from the perspective of complementarity, using TransH, topological relations, and semantic space to calculate loss respectively. By adopting unique learning strategies, we gradually accumulate multimodal knowledge in specific fields and predict answers. On two knowledge-based datasets: ok-vqa and kr-vqa, the performance is higher than that of the latest method respectively. The experimental results prove the validity of the proposed model.