Pre-training A Prompt Pool for Vision-Language Model
Jun Liu, Yang Gu, Zhaohua Yang, Shuai Guo, Huaqiu Liu, Yiqiang Chen · 2023
Pre-trained vision-language(VL) model improves the performance in vision-language tasks, but requires a large amount of labeled data to apply the model to downstream tasks. It is challenging to obtain good results with limited training data. Prompt Tuning(PT), which freezes pre-train language models(PLMs) and only tunes soft prompts, provides an effective solution for adapting PLMs to downstream tasks. However, PT performs comparably with Full-model Tuning(FT) when the data are sufficient and performs much worse in few-shot settings, primarily due to the initialization of soft prompts. In this paper, we propose a new training framework Pre-train a Prompt Pool for Vision-Language Models called “P3VLM”, to give better initialization to PT. The objective of P3VLM is to optimize prompts selected by the visual feature, allowing the prompt to learn relevant knowledge in different data domains and store it in the prompt pool. Then we use diverse prompts in the prompt pool as the initialization of PT instead of a single pre-trained prompt. In the pre-train stage, the training samples are used to train a prompt pool by selecting prompts that best match the visual features. In the downstream datasets, the model selects the best matching prompts as the initialization for PT, freezes the PLM, and only tunes the prompts. Extensive experiments show that our method obtains strong performance on two image caption datasets in both zero-shot and few-shot scenarios.