ZPVQA: Visual Question Answering of Images Based on Zero-Shot Prompt Learning

Naihao Hu, Xiaodan Zhang, Qiyuan Zhang, Wei Huo, Shaojie You · IEEE Access · 2025

In recent years, the use of zero-shot learning to solve visual question-answering (VQA) problems has become a common strategy to address the challenges of complex interactions between visual and verbal modalities. Despite the significant progress of large-scale language models (LLMs) in language tasks, their application to visual question-answering tasks is still challenging because of the differences between visual and textual data. To alleviate this problem, we propose the zero-shot prompt learning VQA (ZPVQA) model, which mitigates the differences between visual and textual data through a prompt-based reasoning strategy and reduces the dependence on end-to-end training. The model devises a method of prompt learning by means of designed prompts that enable LLMs to generate caption prompts based on images and then combines the images with the generated caption prompts in order for the LLMs to generate questions and answers related to the images. In this study, we tested the performance of the ZPVQA model on multiple datasets, and achieve a performance improvement of 3.4% on the VQAv2 dataset and 2.6% on the OK-VQA dataset. The experimental results demonstrate that the prompt learning mechanism designed in this study improves the performance of the model in handling complex multi-modal tasks.

Read the paper · More papers on PaperTik