ICQ-TransE: LLM-Enhanced Image-Caption-Question Translating Embeddings for Knowledge-Based Visual Question Answering
Heng Liu, Boyue Wang, Xiaoyan Li, Yanfeng Sun, Yongli Hu, Baocai Yin · IEEE Transactions on Artificial Intelligence · 2025
In knowledge-based visual question answering (KB-VQA), the answer can be naturally represented by translating visual object embedding referred by the question according to the cross-modality relation embedding related to both the question and the image. Though the triplet representation of cross-modality knowledge is plausible and proven effective, these methods often encounter two challenges: 1) The semantic gap between the image and the question makes it difficult to accurately embed the cross-modality relation; 2) The visual objects in the question often have ambiguous references in the input image. To solve the above challenges, we propose the Image-Caption-Question Translating Embeddings (ICQ-TransE), which more effectively models both the cross-modality relation and the head entity of visual objects. Specifically, for cross-modality relation embedding, the designed image-caption-question information transmission mechanism transmits the information flow from image to question through the caption bridge, where the caption simultaneously has the visual content and the textual form. With this powerful bridge, cross-modality information can be more effectively fused, resulting in more precisely encoded relation embeddings. For the visual object embedding, instead of using a fixed number of visual regions as the previous methods, the most relevant visual regions to the question are dynamically selected. Experimental results on OK-VQA and KRVQA challenging datasets verify the effectiveness of ICQ-TransE compared to multiple state-of-the-art methods for visual question answering with knowledge. Our code will be available at https://github.com/cmcv2022/ICQ-TransE.