Probing Multi-modal Machine Translation with Pre-trained Language Model
Yawei Kong, Kai Fan · 2021
Multi-modal machine translation (MMT) aimed at using images to help disambiguate the target during translation and improving robustness, but some recent works showed that the contribution of visual features is either negligible or incremental.In this paper, we show that incorporating pre-trained (vision) language model (VLP) on the source side can improve the multi-modal translation quality significantly.Motivated by BERT, VLP aims to learn better cross-modal representations that improve target sequence generation.We simply adapt BERT to a cross-modal domain for the vision language pre-training, and the downstream multi-modal machine translation can substantially benefit from the pre-training.We also introduce an attention based modality loss to promote the image-text alignment in the latent semantic space.Ablation study verifies that it is effective in further improving the translation quality.Our experiments on the widely used Multi-30K dataset show increased BLEU score up to 6.2 points compared with the text-only model, achieving the state-of-the-art results with a large margin in the semi-unconstrained scenario and indicating a possible direction to rejuvenate the multi-modal machine translation.