Advancements in Knowledge-Based Visual Question Answering Using Large Language Models: A Review

Noorbhasha Junnu Babu, S. P. Rajamohana · 2025

Visual Question Answering (VQA) is a rapidly evolving domain in artificial intelligence, bridging computer vision and natural language processing to enable machines to comprehend and respond to image-based queries. Traditional VQA models primarily rely on visual and textual features, often struggling with questions that demand external knowledge and reasoning. Recent advancements in Knowledge-Based VQA (KB-VQA) address this limitation by integrating explicit knowledge from structured sources like knowledge graphs and implicit knowledge from large language models (LLMs). This paper provides a comprehensive review of state-of-the-art KB-VQA approaches, including Knowledge-Aware Transformers (KAT), Region-based Visual Instruction Tuning (REVIVE), and Heuristic-Prompted Large Language Models (Prophet). These methods enhance reasoning capabilities by fusing multimodal information with external knowledge, significantly improving accuracy on challenging datasets. However, key challenges persist, including effective knowledge alignment, scalability, computational efficiency, and interpretability. This review critically examines these challenges and explores potential directions for future research, emphasizing the need for more sophisticated architectures and efficient knowledge retrieval mechanisms to advance human-like multimodal AI systems.

Read the paper · More papers on PaperTik