Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation

Yaoming Zhu, Zewei Sun, Shanbo Cheng, Luyang Huang, Liwei Wu, Mingxuan Wang · 2023

Recent work has questioned the necessity of visual information in Multimodal Machine Translation (MMT).This paper tries to answer this question and build a new benchmark in this work.As the available dataset is simple and the text input is self-sufficient, we introduce a challenging dataset called EMMT, whose testset is deliberately designed to ensure ambiguity.More importantly, we study this problem in a real-word scenario towards making the most of multimodal training data.We propose a new framework 2 /3-Triplet which can naturally make full use of large-scale image-text and parallel text-only data.Extensive experiments show that visual information is highly crucial in EMMT.The proposed 2 /3-Triplet outperforms the strong text-only competitor by 3.8 BLEU score, and even bypasses a commercial translation system. 1

Read the paper · More papers on PaperTik