Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts

Deepanway Ghosal, Navonil Majumder, Roy Lee, Rada F. Mihalcea, Soujanya Poria · 2023

Visual question answering (VQA) is the task of answering questions about an image.The task assumes an understanding of both the image and the question to provide a natural language answer.VQA has gained popularity in recent years due to its potential applications in a wide range of fields, including robotics, education, and healthcare.In this paper, we focus on knowledge-augmented VQA, where answering the question requires commonsense knowledge, world knowledge, and reasoning about ideas and concepts not present in the image.We propose a multimodal framework that uses language guidance (LG) in the form of rationales, image captions, scene graphs, etc to answer questions more accurately.We benchmark our method on the multi-choice question-answering task of the A-OKVQA, Science-QA, VSR, and IconQA datasets using CLIP and BLIP models.We show that the use of language guidance is a simple but powerful and effective strategy for visual question answering.Our language guidance improves the performance of CLIP by 7.6% and BLIP-2 by 4.8% in the challenging A-OKVQA dataset.We also observe consistent improvement in performance on the Science-QA, VSR, and IconQA datasets when using the proposed language guidances.

Read the paper · More papers on PaperTik