Zero-Shot Cross-Lingual Knowledge Transfer in VQA via Multimodal Distillation
Yu Weng, Jun Dong, Wenbin He, Chaomurilige Chaomurilige, Xuan Liu, Zheng Liu, Honghao Gao · IEEE Transactions on Computational Social Systems · 2024
As multilingual artificial intelligence systems proliferate, achieving robust cross-lingual understanding remains an open challenge. Recent works have made progress on visual question answering (VQA) models by pretraining on large English image-text datasets. However, there is a language gap as most models are English-centric. Existing attempts at multilingual VQA rely on machine translation or multilingual model pretraining, but cannot effectively transfer rich cross-modal knowledge from English models. In this work, we propose the cross-lingual multimodal knowledge transfer (CMKT) framework to efficiently extend English VQA models to non-English languages via knowledge distillation. Specifically, we introduce a code-mixed cross-lingual mask modeling (CCM) method to establish representations for new languages using small image-text data. We also design a multimodal knowledge distillation (MMKD) method to transfer modal understanding from English models by imitating their sequence processing. Experiments on the xGQA benchmark demonstrate that CMKT can effectively improve zero-shot learning and few-shot learning in non-English languages. Our method reduces the data and computation needed to train multilingual VQA models from scratch. The knowledge transfer paradigm enables non-English languages to inherit and generalize the intricate visual-semantic relationships learned from English. The results also show that the proposed method outperforms previous state-of-the-art methods in the zero-shot setting on the xGQA dataset.