Task-Attentive Transformer Architecture for Continual Learning of Vision-and-Language Tasks Using Knowledge Distillation
Yuliang Cai, Jesse Thomason, Mohammad Ghomi Rostami · 2023
The size and the computational load of finetuning large-scale pre-trained neural networks are becoming two major obstacles in adopting machine learning in many applications.Continual learning (CL) can serve as a remedy through enabling knowledge-transfer across sequentially arriving tasks.However, existing CL algorithms primarily consider learning unimodal vision-only or language-only tasks.We develop a transformer-based CL architecture for learning multimodal vision-and-language (VaL) tasks based on dynamic model expansion and knowledge distillation.Additional parameters are used to specialize the network for each task.Our approach, Task Attentive Multimodal Continual Learning (TAM-CL), enables sharing information between the tasks while addressing catastrophic forgetting.Our approach is scalable, requiring little memory and time overhead.TAM-CL reaches SOTA performance on challenging multimodal tasks.