Meta-learning For Vision-and-language Cross-lingual Transfer
Hanxu Hu, Frank Keller · 2023
Current pre-trained vision-language models (PVLMs) achieve excellent performance on a range of multi-modal datasets.Recent work aims at building multilingual versions of such models, and a range of multilingual multimodal datasets have been introduced for this purpose.However, current PVLMs typically perform poorly on such datasets when used for zero-shot or few-shot cross-lingual transfer, especially for low-resource languages.To alleviate this problem, we propose a novel meta-learning fine-tuning framework.Our framework makes it possible to rapidly adapt PVLMs to new languages by using Modelagnostic Meta-learning (MAML) in a novel cross-lingual multi-modal manner.Experiments show that this new method boosts the performance of current PVLMs in both zero-shot and few-shot settings on four different visionlanguage tasks across 14 languages.