PLPP: Prompt Learning with Perplexity is Self-Distillation for Vision-Language Models

Biao Liu, Wenyi Fang, Xiaoyu Wu, Yang Zheng, Hu Zheng, Bo Yuan · 2026

Pre-trained Vision-Language (VL) models such as CLIP have demonstrated their excellent performance across numerous downstream tasks. A recent method, Context Optimization (CoOp), further improves the performance of VL models on downstream tasks by introducing prompt learning. CoOp optimizes a set of learnable vectors, aka prompt, and freezes the whole CLIP model. However, relying solely on CLIP loss to fine-tune prompts can lead to models that are prone to overfitting on downstream tasks. To address this issue, we propose a plug-in prompt-regularization method called PLPP (Prompt Learning with PerPlexity), which uses perplexity loss to regularize prompt learning. PLPP designs a two-step operation to compute the perplexity for prompts: (a) calculating cosine similarity between the weights of the embedding layer and prompts to get labels, (b) introducing a language model (LM) head that requires no training behind text encoder to output word probability distribution. Meanwhile, we unveil that the essence of PLPP is inherently a form of self-distillation. To further prevent overfitting and reduce extra computation introduced by PLPP, we convert hard labels to soft labels and select top-k values to calculate perplexity loss. For accelerating model convergence, we introduce mutual self-distillation learning, that is perplexity and inverted perplexity losses. Experiments on three classification tasks show PLPP performs better than existing methods.

Read the paper · More papers on PaperTik