ViT-Enhanced Prompts: Integrating Pre-Trained Knowledge for Robust Continuous Learning
Xiaoyu Du, Guoqiang Xiao, Michael S. Lew, Song Wu · 2025
Continuous Learning (CL) enables deep neural networks to adapt to evolving data while preserving performance on previously learned tasks, without access to all historical data. A significant challenge in Continuous Learning is the risk of over-fitting to limited new-class data and catastrophic forgetting of prior knowledge. Although recent methods have demonstrated that pre-trained Vision Transformers (ViTs) with prompt tuning can mitigate these issues, these methods rely on fixed, learnable prompts, which limit knowledge transfer and lead to sub-optimal performance. To address these limitations, we propose a ViT-Enhanced Prompts (VEP) framework for the tasks of continuous learning based on the pre-trained ViTs. Our VEP extracts highly distinctive global and local feature representations from the Multi-Head Self-Attention (MSA) and Multi-Layer Perceptron (MLP) modules of ViT, integrating them with dynamic prompts to enhance knowledge learning and transferring. Additionally, we introduce a diversity loss mechanism to reduce redundant information accumulation in the learned prompts and enhance their discrimination, aiming to mitigate catastrophic forgetting effectively. Extensive experiments on benchmark datasets, including ImageNet-R, CUB200, and CIFAR100, validate the contribution of each component in VEP and demonstrate that VEP achieves state-of-the-art performance on the task of Continuous Learning while maintaining low parameter overhead. The source code of our designed VEP is at https://github.com/SWU-CSMediaLab/VEP.