Unveiling the power of transfer learning towards efficient artificial intelligence

Can Qin · 2023

Large-scale models, abundant data, and dense computation are the pivotal pillars of deep neural networks. The present-day deep learning models have made significant strides in various areas such as Computer Vision (CV), Natural Language Processing (NLP), and Audio Signal Processing (ASP). These technological integrations have notably improved industrial automation while providing considerable enhancements to daily life. However, despite these advancements, deep learning still faces severe challenges in evolving into an efficient and accessible system. One of the major concerns is data efficiency due to the labor-intensive and costly process of annotated data. The other concern is model efficiency, impacting deployment costs and users' accessibility. Transfer Learning (TL) is a promising solution to address these challenges. TL harnesses the power of acquired data and pre-trained models to facilitate applications of new related tasks or smaller models. This dissertation is structured into three primary sections: Feature Transfer Learning, Model Transfer Learning, and Joint Transfer Learning. (1) Feature Transfer Learning (FTL), widely employed in Domain Adaptation (DA), utilizes a shared encoder model to learn universal representations through cross-domain feature alignment loss. It is primarily comprised of Unsupervised Domain Adaptation (UDA) and Semi-supervised Domain Adaptation (SSDA), depending on target label accessibility. The principal technical challenges with FTL involve distribution mismatch across domains and overfitting toward labeled data. To address these issues, this dissertation proposes structural regularization and multi-level alignment. (2) Model Transfer Learning (MTL) focuses on parameter tuning based on pre-trained models for novel tasks. An exemplary application of MTL is Knowledge Distillation (KD), which facilitates knowledge transfer from larger to smaller models for compression. This dissertation introduces a graph-based KD framework that enables real-time graph retrieval. In addition, with the surge of foundation models necessitating efficiency during finetuning, Parameter-Efficient Model Finetuning (PEFT) has received prominence. PEFT has been applied here to enrich a pre-trained tabular model's capacity by injecting external prior knowledge. (3) Joint Transfer Learning (JTL) synergizes FTL and MTL, necessitating both cross-domain feature alignments and parameter tuning. JTL is particularly suitable for instance alignment across different modalities, which helps to build multimodal models without --Author's abstract

Read the paper · More papers on PaperTik