A Language-Driven Data Augmentation Method for Mongolian-Chinese Neural Machine Translation

Xuerong Wei, Qing-Dao-Er-Ji Ren · 2024

Traditional data augmentation methods are typically computation-driven, randomly selecting words for modification with equal probability for each word. However, these methods do not take into account the linguistic information conveyed between words, which can disrupt the grammatical structure of the sentence and reduce text quality. In this paper, a language-driven data augmentation method for Mongolian-Chinese neural machine translation is proposed to address these problems. Specifically, the Stanford CoreNLP is first used to construct the dependency tree of the Chinese sentences. Next, the PageRank algorithm is employed to calculate the importance of each word in the sentence, and standard word-level data augmentation is performed on words with lower importance to generate new sentences. These new sentences maintain the same Mongolian alignment as the original sentences. Finally, the newly generated parallel corpus is combined with the original parallel corpus to form a pseudo-parallel corpus, which aids in the training of the neural machine translation model. The experimental results on the Mongolian-Chinese parallel corpus presented in this paper show that the data augmentation method proposed in this paper has a significant improvement in BLEU values over the baseline translation model.

Read the paper · More papers on PaperTik