Effective Retrieval Augmentation for Knowledge-Based Vision Question Answering

Jiaqi Deng · 2024

Knowledge-based Vision Question Answering (KBVQA) systems aim to answer natural language questions grounded in image contents by retrieving and integrating relevant knowledge from external knowledge bases to generate the final answers. These systems have diverse application scenarios ranging from general cross-modal understanding to specialized domains like healthcare or robotics. While significant progress has been made, several key obstacles still persist: (1) multimodal pretrained models remain underinvestigated in KB-VQA, and existing methods can incur high computational costs; (2) Current KB-VQA systems tend to extract noisy information from external knowledge sources, which might impair the answer quality; (3) Cross-modal interactions between images and textual questions are still insufficient at the current state, potentially limiting the system's ability to comprehend essential contents. This doctoral consortium paper intends to provide a comprehensive review of existing efforts in KB-VQA systems and offer insights into potential future research directions in this domain. We highlight the importance of improving multimodal pretraining techniques, mitigating noise in knowledge integration, and enhancing interaction modeling between images and texts to improve the performance and the robustness of KB-VQA systems. Proposed research methodology will also be outlined in this paper.

Read the paper · More papers on PaperTik