Progressive Multi-Prompt Learning for Vision-Language Models
Jun Liu, Ziqian Lu, Hao Luo, Zhe‐Ming Lu, Yangming Zheng · IEEE Transactions on Circuits and Systems for Video Technology · 2025
Recently, methods that utilize prompt tuning to rapidly transfer pretrained vision-language models (VLMs) to downstream tasks have been proposed. Although these models have produced reasonable results, they typically learn a single prompt, which limits their ability to capture more diverse information. This ability is crucial for addressing fine-grained classification challenges and intraclass visual variability (e.g., color, pose, and size variations within the same category). However, learning multiple prompts provides a larger optimization space, which further exacerbates the overfitting phenomenon. This makes it more challenging balance the performances acienved for base and new categories. To address these issues, we propose progressive multi-prompt (PMP) learning method.ecently, methods that utilize prompt tuning to rapidly transfer pretrained vision-language models (VLMs) to downstream tasks have been proposed. Although these models have produced reasonable results, they typically learn a single prompt, which limits their ability to capture more diverse information. This ability is crucial for addressing fine-grained classification challenges and intraclass visual variability (e.g., color, pose, and size variations within the same category). However, learning multiple prompts provides a larger optimization space, which further exacerbates the overfitting phenomenon. This makes it more challenging balance the performances acienved for base and new categories. To address these issues, we propose progressive multiprompt (PMP) learning method.R Specifically, we introduce multiple prompts in a step-by-step manner to focus on various information. To reduce overfitting, we utilize alate attachingmechanism to defer the interactions of prompts and features to a deeper encoding layer. Furthermore, we balance the prompts for different layers with learnable weights to guide the optimal optimization procedure. We compared our method with several state-of-the-art approaches in base-to-new task settings and demonstrate superior base-new tradeoff performance. Additionally, we conducted cross-dataset transfer, domain generalization, and few-shot experiments to further validate the effectiveness of our method. Our code is available at https://github.com/JunLGeek/PMP.git.