TransferCVLM: Transferring Cross-Modal Knowledge for Vision-Language Modeling
Dongha Choi, Jung‐Jae Kim, Hyunju Lee · 2024
Recent large vision-language multimodal models pre-trained with huge amount of imagetext pairs show remarkable performances in downstream tasks.However, the multimodal pre-training has limitations in terms of resources and training time when it comes to obtaining new models that surpass existing models.To overcome these issues, we propose TransferCVLM, a method of efficient knowledge transfer that integrates pre-trained uni-modal models (and cross-modal fusionencoder) into a combined vision-language model (CVLM), without pre-training the CVLM with large amount of multimodal data, and then for each task application, fine-tunes the CVLM and transfers the multimodal knowledge of a teacher vision-language model to the CVLM by using knowledge distillation techniques.We demonstrate that 1) the finetuned CVLM performs comparable to other vision-language models of similar size, that 2) the multimodal knowledge transfer consistently enhances the CVLM, and the knowledgetransferred CVLM composed of large-size unimodal models outperforms the teacher multimodal model in most of downstream tasks, and that 3) TransferCVLM can also be used for model compression when using small-size unimodal models.We estimate that the training of TransferCVLM takes only 6% of pretraining of other vision-language models.Our code is available at https://github.com/DMCB-GIST/TransferCVLM.