Parameter-Efficient Cross-lingual Transfer of Vision and Language Models via Translation-based Alignment

Zhen Zhang, Jialu Wang, Xin Long Wang · 2023

Pre-trained vision and language models such as CLIP (Radford et al., 2021) have witnessed remarkable success in connecting images and texts with a primary focus on English texts.Despite recent efforts to extend CLIP to support other languages, disparities in performance among different languages have been observed due to uneven resource availability.Additionally, current cross-lingual transfer methods of those pre-trained models would consume excessive resources for a large number of languages.Therefore, we propose a new parameter-efficient cross-lingual transfer learning framework that utilizes a translationbased alignment method to mitigate multilingual disparities and explores parameterefficient fine-tuning methods for parameterefficient cross-lingual transfer.Extensive experiments on XTD (Aggarwal and Kale, 2020) and Multi30K (Elliott et al., 2016) datasets, covering 11 languages under zero-shot, few-shot, and full-dataset learning scenarios, show that our framework significantly reduces the multilingual disparities among languages and improves cross-lingual transfer results, especially in low-resource scenarios, while only keeping and fine-tuning an extremely small number of parameters compared to the full model (e.g., Our framework only requires 0.16% additional parameters of a full-model for each language in the few-shot learning scenario).The codes are available at https://github.com/ eric-ai-lab/PECTVLM.

Read the paper · More papers on PaperTik