Unsupervised Domain Adaptation for Vision Transformers Based on Prompt Learning
Hong Zhi Yu · 2024
Currently, in the field of deep learning, the “pre-train and fine-tune” paradigm has become the mainstream training method for many network models due to the increased capacity of network models and the variety of tasks. However, fine-tuning has become increasingly challenging, as a new model needs to be fine-tuned for each task, which consumes a lot of time and computational resources. This paper proposes an unsupervised domain-adapted image classification method based on prompt learning and Vision Transformer. Unlike the traditional “pre-training, fine-tuning” paradigm, the method applies the pre-trained models of prompt learning and Vision Transformer to the unsupervised domain adaptation task, and proposes to redefine image classification as a mask prediction task. Specifically, the pre-trained model is guided by constructing templates and labeled word sets for prompt learning. Then, the image classification problem is converted into the form of a complete blank and embedded into the constructed template, and the labeled words corresponding to the maximum probability are selected and mapped back to the set of labeled words by the pre-trained model. By conducting experiments on several UDA benchmark datasets, including 93.5% for Office-31, 83.30% for Office-Home, and 81.97% for VisDA-2017, it can be demonstrated that prompt learning greatly improves the efficiency of using the pre-trained model and saves computational resources without significantly affecting the model performance.