Adaptive Transfer Learning via Fine-grained Multi-task Pre-training
Yanbao Ma, Hao Xu, Jun-Zhou He, Kun Qian, Tiebing Li · 2021 4th International Conference on Algorithms, Computing and Artificial Intelligence · 2021
Nowadays pre-training paradigm has been widely adopted for deep learning-based applications. In multiple pre-training tasks, conventional methods process them using naive Multi-Task Learning (MTL) technology. The pre-trained models are unavoidably influenced by the well-known negative transfer phenomenon of MTL, which will also make a negative impact on downstream tasks. To deal with this problem, we propose a novel Adaptive Pre-Training (APT) framework. In the pre-training stage, we adopt a task-sensitive MTL technology called Fine-grained Sharing Network (FSN), which trains a subnet mask for each task besides the network weights. In the fine-tuning stage, the most suitable subnet is selected for each downstream task. Therefore, different downstream tasks may be fine-tuned based on different network structures and benefit from the outcome of their most closely related pre-training task. According to our experiments, the proposed APT outperforms the conventional pre-training baseline at a clear margin. In addition, experiments show that, even without the subnet adaptation, the network weights pre-trained by FSN alone are of much higher quality and can be used in the framework of conventional pre-training to improve performance.