Student Can Also be a Good Teacher
Jun Rao, Qian Tao, Shuhan Qi, Yulin Wu, Qing Min Liao, Xuan Wang · 2021
Astounding results from transformer models with Vision-and Language Pretraining (VLP) on joint vision-and-language downstream tasks have intrigued the multi-modal community. On the one hand, these models are usually so huge that make us more difficult to fine-tune and serve real-time online applications. On the other hand, the compression of the original transformer block will ignore the difference in information between modalities, which leads to the sharp decline of retrieval accuracy.