A Vision-language Model Based on Prompt Learner for Few-shot Medical Images Diagnosis

Tianyou Chang, Shizhan Chen, Guodong Fan, Zhiyong Feng · 2024

In the real world, it can be challenging to annotate a large-scale dataset for all medical images, making few-shot medical image classification an important task. The latest advancements in pre-trained vision-language models as CLIP have demonstrated excellent performance in zero-shot natural image recognition and show advantages in medical applications. However, we have found that deploying such models in practical applications faces challenges in terms of engineering effort. It requires specialized medical domain knowledge and is time-consuming, as even slight variations in wording can have a significant impact on performance. Inspired by recent research on prompt learning in the field of Natural Language Processing (NLP), we propose a simple approach called Prompt Learner (PoLe) for automating the design of prompts in pre-trained vision-language models. This is a simple method specifically designed to fine-tune vision-language models, similar to CLIP, for downstream image recognition. In particular, PoLe models context tokens using continuous vectors that can automatically learn from medical images, thus avoiding the tedious process of handcrafting prompt engineering. Additionally, this approach maintains the frozen state of the large-scale pre-trained parameters, saving computational resources. Through extensive experiments on 5 medical image datasets, we have demonstrated that PoLe surpasses manually designed prompts with just one or two shots, and further training with more shots significantly improves the performance of image classification. For instance, when trained with 16 shots, the average improvement is approximately 20% (with a maximum improvement of over 31%). PoLe effectively transforms CLIP into a powerful few-shot learner. In terms of recognition performance, adjusting the CLIP model using PoLe yields better results than manually designed prompts for CLIP. When enhancing CLIP, PoLe demonstrates stronger learning capabilities compared to other few-shot learners such as linear probes. Furthermore, it outperforms CLIP models assisted by ChatGPT on most datasets. This indicates that PoLe possesses significant adaptability in the field of medical image analysis.

Read the paper · More papers on PaperTik