DueT: Image-Text Contrastive Transfer Learning with Dual-adapter Tuning
Taku Hasegawa, Kyosuke Nishida, Koki Maeda, Kuniko Saito · 2023
This paper presents DueT, a novel transfer learning method for vision and language models built by contrastive learning.In DueT, adapters are inserted into the image and text encoders, which have been initialized using models pre-trained on uni-modal corpora and then frozen.By training only these adapters, DueT enables efficient learning with a reduced number of trainable parameters.Moreover, unlike traditional adapters, those in DueT are equipped with a gating mechanism, enabling effective transfer and connection of knowledge acquired from pre-trained uni-modal encoders while preventing catastrophic forgetting.We report that DueT outperformed simple finetuning, the conventional method fixing only the image encoder and training only the text encoder, and the LoRA-based adapter method in accuracy and parameter efficiency for 0-shot image and text retrieval in both English and Japanese domains.