Policy Gradient-Based Optimal Subset Selection for Few-Shot Vision-Language Learning

Muhammad Khizer Ali, Manoranjan Paul, Anwaar Ulhaq, Muhammad Haris Khan, Quazi Mamun · 2025

Vision-Language models (VLMs) like Contrastive Language-Image Pre-Training (CLIP) have been extensively adapted for few-shot classification. Most few-shot methods rely on randomly selected samples from the dataset. However, since only a few samples are used, the sample selection process can significantly impact the performance of the downstream classification task. In this work, we propose a reinforcement learning-based policy gradient technique that employs a diversity and informativeness-based reward function to optimise the sample selection process. We evaluate various sample selection techniques based on downstream classification accuracy across three benchmark datasets, where the proposed method demonstrates promising results.

Read the paper · More papers on PaperTik