Cross-language Text Generation Using mBERT and XLM-R: English-Chinese Translation Task
Ying Ma · 2024
Traditional methods often require a large number of parallel corpora, that is, texts from one language and corresponding texts from another language. The acquisition and construction of this data is very expensive and time-consuming, which limits the application scope and feasibility of the model. This article used mBERT and XLM-R to solve the problem of relying on parallel corpora in traditional methods, and improved the quality of English-Chinese translation through the model's own multilingual representation ability. A publicly available parallel corpus of English-Chinese was collected for translation tasks, and the collected corpus was preprocessed with segmentation, punctuation, and other techniques. mBERT and XLM-R were used for English-Chinese translation respectively, and Ensemble method was adopted to weight and fuse the results of mBERT and XLM-R. BLEU (Bilingual Evaluation Understudy), METEOR (Metric for Evaluation of Translation with Explicit Ordering), and ROUGE (Recall-Oriented Understudy for Gisting Evaluation) were used to evaluate the quality of the translations. The average BLEU, average METEOR, and average ROUGE-L values of mBERT-XLM-R were 0.98, 0.97, and 0.98, respectively, according to the results. Therefore, mBERT and XLM-R working together may significantly enhance the quality of English-Chinese translation.